Skip to content
Article

What Is an LLM (Large Language Model) and How Does It Work?

A large language model is an AI model trained on huge text corpora to generate text. Tokens, training, GPU needs, limits and security risks explained.

Oğuzhan Gerçek··7 min read
What Is an LLM (Large Language Model) and How Does It Work?

Short answer: An LLM (large language model) is a deep learning model pre-trained on very large text corpora that can understand and generate text. At its core it predicts which token (a word or part of a word) comes next in a piece of text; by making that prediction over and over, it produces answers, summaries, translations or code. Today's LLMs are built on the Transformer architecture, published in 2017. They are powerful but not infallible: they can state wrong information with confidence, and they know nothing about events after their training data ends.

What is an LLM?

In AWS's definition, LLMs are very large deep learning models that are pre-trained on vast amounts of data. "Large" refers to the number of parameters: the weights, biases and embeddings that the model adjusts during training. For a sense of scale, the GPT-3 model described in OpenAI's 2020 paper has 175 billion parameters.

An LLM is the branch of AI that specializes in language: deep learning sits inside machine learning, and these large language-focused models sit inside deep learning. As our glossary entry on LLMs points out, in enterprise use the real question is less the model's capability than which data goes where.

The Transformer: the architecture behind LLMs

Today's LLMs trace their architecture to "Attention Is All You Need", published by Vaswani and colleagues in 2017. The paper proposed a new architecture based solely on attention mechanisms, dropping recurrence and convolutions entirely: the Transformer. As AWS puts it, this structure extracts meaning from a sequence of text and understands the relationships between its words and phrases.

One of the paper's key findings was speed: the Transformer was more parallelizable and took significantly less time to train. An architecture that runs in parallel maps directly onto hardware that runs thousands of threads at once, the GPU.

Tokens, pre-training, fine-tuning and inference

  • Token: A model processes text by breaking it into units called tokens. According to Microsoft Learn, a token can be a word, a set of characters or a combination of words and punctuation, and the tokenization method varies by model. The combined limit on input and output tokens is called the context window, and the cost of each request depends on the number of tokens.
  • Pre-training: On a large text corpus, the model adjusts its parameters over and over until it correctly predicts the next token from the previous ones.
  • Fine-tuning: A pre-trained base model is trained further on a relatively small, task-specific dataset to adapt it to a particular job.
  • Inference: The trained model producing a response to new input; every user question is an inference request.

Token counts also vary by language. A study published at NeurIPS 2023 found that the same text translated into different languages can split into very different numbers of tokens, with differences of up to 15 times in some cases. If your content is not in English, measure cost and context window on your own texts rather than on English examples.

Why do LLMs need so many GPUs?

The answer is memory. By NVIDIA's calculation, a 7-billion-parameter model loaded in 16-bit precision needs roughly 14 GB of memory for its weights alone. By the same math, the weights of a 70-billion-parameter model come to 140 GB, which will not fit on an H100 GPU with 80 GB of memory. Inference also needs memory beyond the weights for the so-called KV cache, which, according to NVIDIA, grows linearly with the number of requests processed together and with sequence length.

One way around this is to distribute the model over several GPUs, reducing the memory load on each. We discuss the choice between buying and renting that hardware in our article on buying vs renting GPUs.

Limits of LLMs

  • Hallucination: In its generative AI profile (AI 600-1), NIST calls this "confabulation": the production of confidently stated but erroneous or false content. NIST describes it as a natural result of how generative models are designed, since they produce outputs that approximate the statistical distribution of their training data.
  • Stale knowledge: As AWS's RAG page puts it, LLM training data is static and introduces a cut-off date on the model's knowledge. When a user expects a specific, current answer, the model may respond with out-of-date or generic information.
  • Cost: Every request consumes tokens, so the bill grows with usage; we cover managing inference costs in our AI FinOps article.

Enterprise use: RAG, open models and data privacy

In enterprise use, a model also has to answer from knowledge that is not in its training data: internal policies, product documentation, contracts. One way to do this is RAG (retrieval-augmented generation). In AWS's description, RAG has the model reference an authoritative knowledge base outside its training data before generating a response, so it can draw on an organization's internal knowledge without being retrained and can cite its sources.

The second decision is where the model runs:

  • A model consumed through an API: You can start quickly without building infrastructure, but prompts and the data in them go to the provider's systems.
  • An open-weight model on your own infrastructure: The data stays with you, but GPU capacity, model versions, evaluation and on-call duty become your job.

What decides which workload goes where is not whether the model is open or closed but the class of the data; we explain this in our sovereign AI article. If prompts contain personal data, the country where the model runs matters too. We cover what that means under the KVKK (Türkiye's personal data protection law) in our article on KVKK and cross-border transfers, and how to govern AI use inside the organization in our AI governance article.

Security risks: the OWASP LLM Top 10

The OWASP GenAI Security Project collects the most critical security risks for LLM applications in its community-driven Top 10 for LLM Applications. The current edition is the 2026 list, and its top three are:

  1. Prompt injection (LLM01): Any input to the model, whether user input, a document the model reads or a tool's output, altering the model's behavior in ways the developer did not intend. The instruction may be hidden in content the user neither wrote nor saw.
  2. Sensitive information disclosure (LLM02): An LLM-integrated system exposing confidential or regulated information.
  3. Excessive agency (LLM03): Giving an LLM that can call tools more functionality, permissions or autonomy than it needs. It was sixth on the 2025 list; the project team calls its climb the most consequential move on the list.

The list also includes unbounded consumption (LLM06) and misinformation (LLM07), which covers hallucination. We discuss the order in which authority should be handed to agents in our article on AI agents in IT operations, and how we design this layer on our AI stack design and engineering page.

Frequently asked questions

What does LLM stand for? Large language model: an AI model trained on very large text corpora that can understand and generate text.

What are LLMs used for? Language tasks such as writing copy, answering questions from knowledge bases, classifying text and generating code.

What is the difference between an LLM and AI? AI is the broadest concept; an LLM is one kind of deep learning model within it, specialized in language. Every LLM is an AI system, but not every AI system is an LLM.

How is an LLM trained? It is first pre-trained on a very large text corpus to predict the next token. It can then be fine-tuned for specific tasks on smaller, task-specific data.

What is an LLM hallucination? The model confidently producing incorrect or false information. NIST calls this risk "confabulation".

What is RAG? Retrieval-augmented generation: the model consults an external knowledge source, such as company documents, before generating a response. It lets the model answer with current, organization-specific information without retraining.

Sources