Skip to content
Article

What Is a GPU? GPU vs CPU and the Role of GPUs in AI

A GPU is a processor built to handle large amounts of data at once. Why it works in parallel, training vs inference, and how data center GPUs differ.

Oğuzhan Gerçek··7 min read
What Is a GPU? GPU vs CPU and the Role of GPUs in AI

Short answer: A GPU (graphics processing unit) is a processor designed to work on large amounts of data at the same time. It was first developed to speed up graphics rendering; today it is used for machine learning, video editing and gaming. A CPU uses a small number of powerful cores to finish a sequence of operations as fast as possible, while a GPU runs thousands of threads at once. Because AI models are built from layer upon layer of linear algebra, the GPU has become the key hardware for deep learning.

What is a GPU?

Intel defines the GPU as a specialized processor originally designed to accelerate graphics rendering. In graphics, textures, lighting and the rendering of shapes all have to be done at once. According to NVIDIA, what makes the GPU suited to this is its ability to break a complex problem into thousands or millions of separate tasks and work them out simultaneously. The same ability later took it beyond graphics.

There are two kinds: an integrated GPU does not come on a separate card but is embedded alongside the CPU, while a discrete GPU is a separate chip mounted on its own circuit board. The terms GPU and graphics card are often used interchangeably, but Intel draws a line between them: just as a motherboard carries the CPU, a graphics card is the add-in board that carries the GPU.

How does a GPU differ from a CPU?

NVIDIA's CUDA programming guide explains the difference through design goals. A CPU is designed to execute a sequence of operations, called a thread, as fast as possible, and can run a few tens of these threads in parallel. A GPU runs thousands of them at once, accepting slower single-thread performance in exchange for greater overall throughput. According to the guide, a GPU devotes more transistors to data processing rather than to data caching and flow control. NVIDIA's article on CPUs and GPUs sums up the same split in terms of cores: a CPU has just a few cores with lots of cache memory, while a GPU has hundreds of cores that can handle thousands of threads simultaneously.

The two are not rivals. Because an application usually mixes parallel and sequential parts, systems are designed to use GPUs and CPUs together to maximize overall performance.

From graphics to general-purpose computing: CUDA

In November 2006, NVIDIA introduced CUDA, a general-purpose parallel computing platform and programming model that uses the parallel compute engine in NVIDIA GPUs to solve complex computational problems. According to a 2009 NVIDIA article, the platform was first released in 2007 and let developers put the GPU's computing power to work on general-purpose processing.

CUDA was built for NVIDIA GPUs; for AMD GPUs there is ROCm, AMD's open software stack for GPU-accelerated computing. Which platform your AI libraries support also narrows your hardware options, so choosing a GPU is a software stack decision as much as a hardware one.

Why does AI run on GPUs?

According to NVIDIA's article on GPUs and AI, AI models are made of layer upon layer of linear algebra equations, and a GPU works through that math in parallel across thousands of cores. Recent GPUs also include Tensor Cores dedicated to the matrix math that neural networks use. Training and running a deep learning model means repeating these operations at a very large scale. We look at the GPU demands of large language models separately in our LLM article.

Training vs inference

GPUs serve two different jobs in AI:

  • Training: The model learns from data. In NVIDIA's explanation, the model's error is propagated back through the network's layers, and the model adjusts its weights and tries again. The process consumes an enormous amount of compute.
  • Inference: The trained model is applied to new, unseen data and returns answers quickly and accurately. Every question put to an AI assistant is an inference request.

Training is intensive and periodic; once the model is ready, the capacity frees up. Inference continues for as long as the model stays in production and grows with the number of users. We look at how to manage the inference bill in our AI FinOps article.

Data center GPUs vs consumer GPUs

A gaming GPU and a data center GPU may come from the same vendor, but they are designed for different jobs. Examples from NVIDIA's own product pages:

  • Memory: The H200 data center GPU has 141 GB of HBM3e memory with 4.8 TB/s of bandwidth. The consumer GeForce RTX 5090 has 32 GB of GDDR7. This number decides whether a model fits on a single card.
  • ECC and 24/7 operation: NVIDIA lists ECC memory for the L40S and says the card is optimized for 24/7 enterprise data center operations. We explain what ECC does in our server article.
  • Cooling: Cards such as the L40S and H100 NVL are passively cooled. According to the H100 NVL product brief, the passive heat sink requires system airflow to keep the card within its thermal limits; the card relies on the server's cooling, not on fans of its own.
  • Power and sharing: The SXM version of the H200 has a power limit of up to 700 W; we explain why per-rack power and cooling must be calculated separately in our data center article. The same GPU can be split into up to seven instances with MIG (Multi-Instance GPU) and shared across several workloads.
  • Licensing: Under NVIDIA's GeForce software license, GeForce and Titan drivers are licensed only for use on GeForce or Titan hardware you own, and the license does not cover data center deployment.

A gaming card can do the job for experiments and development. For a service that will run continuously in a data center, memory, cooling and licensing each need their own answer.

Should you buy or rent GPUs?

Once the need for GPUs is clear, the next question is whether to buy the hardware or rent it. Two things largely decide it: how much of the time the GPUs will be busy, and where the data is allowed to be processed. We cover how to calculate the utilization threshold and how data residency changes the math in our article on buying vs renting GPUs, and the budget side in our CAPEX and OPEX article. Our sovereign AI article explains why separating workloads that handle sensitive data is the first step.

If you choose to own, a GPU server cannot be planned like an ordinary server: power, cooling, the network between GPUs and sharing between teams must be designed in from the start. We describe how we design this layer on our AI stack design and engineering page.

Frequently asked questions

What does GPU mean? GPU stands for graphics processing unit: a processor designed to work on large amounts of data at the same time.

What is a GPU used for? It was first developed to speed up graphics rendering. Today, besides gaming, it is used for machine learning, for training and running AI models, and for video editing.

What is the difference between a GPU and a CPU? A CPU runs a sequence of operations as fast as possible on a small number of powerful cores. A GPU runs thousands of threads in parallel, which is why it pulls ahead on work that repeats the same operation over large amounts of data.

Is a graphics card the same thing as a GPU? The terms are often used interchangeably, but they are not quite the same. The GPU is the processor itself; the graphics card is the board that carries it and plugs into the computer.

Can a gaming graphics card be used for AI? Yes, for experiments and development with small models. But it has less memory than data center cards, and NVIDIA's GeForce driver license does not cover data center deployment.

Sources