myibrahim.cloud

AI / ML · Foundations

LLMs, explained for engineers who use them daily

What's actually happening inside that API call. Tokens, attention, context windows, sampling — the parts that change how you write code with them.

You don't need a PhD to use LLMs well. You don't need to read the Attention Is All You Need paper. You do need to know roughly what's happening when you call claude.messages.create(...), because the abstractions leak in ways that affect your bill, your latency, and your prompt design.

This is the version I wish someone had handed me eighteen months ago.

Karpathy: Intro to Large Language Models (1hr)

What an LLM actually does#

An LLM is a function. The argument is a sequence of tokens. The return value is a probability distribution over the next token. That's it.

Everything else — the chat interface, the streaming, the system prompts, the tool use — is glue around that one operation, called over and over until something says stop.

Tokens, not words#

The model doesn't see English. It sees tokens. Tokens are not words and not characters; they're sub-word chunks the tokenizer chose by statistics. The word "tokenization" might split into ["token", "ization"]. The word "Mohammed" might be one token or three. Punctuation and whitespace become tokens too.

Two practical consequences:

  • Cost is per token, not per word. A long English word is ~1.3 tokens. Code is ~0.4 tokens per character. Languages with rare scripts (Arabic, Korean, Tamil) cost more per "word" than English does because their tokenizers split more aggressively. Bills are skewed toward verbose-language users.
  • Context limits are in tokens. A 200K-token context window holds ~150K English words, less for code, much less for non-English.

Here's how you'd inspect it in Python with Anthropic's tokenizer (or tiktoken for OpenAI):

from anthropic import Anthropic
client = Anthropic()

# Quick estimate
text = "Tokenization is the secret bottleneck."
print(client.count_tokens(text))  # ~9 tokens

When prompts get long, count tokens before you call. Surprises here are expensive.

The context window is a sliding scratchpad#

The model has no memory between calls. Whatever you pass in messages=[...] is its entire universe. If a previous turn wasn't in that array, it doesn't exist.

This is why "long conversations" require you (the engineer) to maintain history and re-send it. The model is stateless. The chat UIs you've used add the state on top.

Implication: as a conversation grows, every turn re-pays for the entire history. A 10-turn chat where each turn adds 500 tokens has token bill ~`10 + 9 + 8 + ... = 55× turn cost, not 10×`. Caching helps a lot here.

Tokens billed per turn (uncached)
  • Turn 1500
  • Turn 52,500
  • Turn 105,000
  • Turn 2010,000

Attention: why the order matters#

Attention is the mechanism by which each token "looks at" the others when forming its prediction. The headline simplification: every token weights every other token in its context, asks "how relevant are you to what I'm trying to say next," and pools their representations weighted by relevance.

Three things follow from this for you:

  • Position matters. The model has been trained on positional patterns. The same instruction at the top vs. middle vs. bottom of a long context produces different completions. There's a phenomenon called "lost in the middle": models recall content near the start and end of a long context better than the middle.
  • Long context degrades. Throwing 150K tokens at a 200K-window model is technically allowed, but model performance on retrieval tasks drops noticeably past ~30K. Don't assume a bigger window is a substitute for retrieval.
  • Format leaks. Markdown, XML, JSON — the model picks up on the structural cues. A prompt that says "the user message is below" with the message in <user>...</user> tags works better than the same content with no boundaries.

Sampling: temperature, top-p, top-k#

The model gives a probability distribution. Sampling chooses one token from it.

  • temperature=0 (greedy): always pick the most likely. Deterministic, but in practice modern LLMs still aren't 100% deterministic at temp 0 due to floating-point non-determinism. Closer than not.
  • temperature=1.0: stock distribution.
  • temperature=2.0: the distribution gets flatter; rarer tokens become viable. Output gets weirder.

Top-p (nucleus): "only consider tokens that together account for the top p% of probability mass." top_p=0.9 is a sane default for most chat use.

Top-k: "only consider the k most likely tokens." Less common in modern APIs.

For code generation: temperature=0 or 0.2. For brainstorming: 0.7–1.0. For "creative writing": maybe 1.0+ with top_p=0.9. For evaluation runs where you want reproducibility, set both.

RLHF and instruction tuning#

A raw "base" LLM trained on internet text is good at one thing: continuing internet text. Ask it a question and you might get back a list of related Reddit posts. Helpful as an autocomplete; useless as an assistant.

The models you use have been fine-tuned with two main steps after pre-training:

  1. Supervised fine-tuning (SFT) on human-written demonstrations of "good answers."
  2. Reinforcement learning from human feedback (RLHF) — humans rank outputs, a reward model is trained on those rankings, the LLM is fine-tuned to maximize the reward.

This is why modern LLMs sound so eager to please. It's also why they hedge, refuse, and over-explain. The reward model rewarded those behaviors during training. They are policy decisions baked into the weights.

When a model refuses something it shouldn't, or hedges where you wanted a direct answer, the rewards-during-training is the thing you're fighting. Sometimes a different model trained with a different rewards mix behaves better on your task. The right model isn't always the biggest.

Function calling / tool use is just structured output#

The "agent" frameworks make tool use sound exotic. Underneath, it's the same single-token-at-a-time generation, except the model has been fine-tuned to emit structured JSON when it wants the orchestration layer to do something on its behalf, and to expect the result of that something to be added back to the conversation.

result = client.messages.create(
    model="claude-opus-4-7",
    tools=[{
        "name": "get_weather",
        "description": "Look up weather by city.",
        "input_schema": {
            "type": "object",
            "properties": {"city": {"type": "string"}},
            "required": ["city"],
        },
    }],
    messages=[{"role": "user", "content": "What's the weather in Cairo?"}],
)

if result.stop_reason == "tool_use":
    tool_call = next(b for b in result.content if b.type == "tool_use")
    weather = my_weather_api(tool_call.input["city"])
    # Append the tool result to messages and call again, model generates final answer

The model isn't calling the tool. It's emitting tokens that say "I'd like this tool called with these arguments." Your code does the call and feeds the result back. The model continues from there.

That's also why a model with more tools available in a single turn often performs worse at any one of them — it has to allocate attention and instructions across more options. Keep tools focused.

Streaming is incremental rendering#

The API can return tokens as they're generated rather than waiting for the full response. This isn't faster overall — the model produces the same total tokens — but the UX feels alive.

with client.messages.stream(...) as stream:
    for text in stream.text_stream:
        print(text, end="", flush=True)

The wins from streaming:

  • Time-to-first-token (TTFT) is what feels like "the model thinking." Often <500ms.
  • Long generations show progress, so users don't bail.
  • You can interrupt mid-stream when the user navigates away.

Cost note: most providers don't bill differently for streamed vs. non-streamed responses. Streaming is purely a UX choice.

What this means for your code#

A few practical takeaways:

  1. Token-aware design. Count tokens before sending; cap context. Truncate old turns or summarize them. Always have a budget.
  2. Prompts as code. Version them. Test them. The same system prompt in production-prompt vs. dev-prompt produces measurably different outputs. Treat it like config.
  3. Cache the static parts. Anthropic's prompt caching lets you mark long, stable preludes (your system prompt, examples, retrieved docs) as cacheable. Cache hits are ~10× cheaper. Pay attention to cache TTL (5 minutes default).
  4. Defensive parsing. Even with structured output, validate the JSON you got. The model is a probabilistic system; assume failures.
  5. Eval before scaling. Don't ship LLM features without an eval set. The gap between "looks good in the demo" and "works at p99 in production" is where most LLM apps die.

The next articles in this section unpack each of these — RAG, embeddings, fine-tuning, eval. Start there if you have a specific feature to ship. Come back here when something behaves weirdly and you need the mental model to debug it.

Further reading#

  • Karpathy, "Intro to Large Language Models" (video, linked above) — the canonical 1hr primer.
  • Anthropic's Building with Claude docs — practical, opinionated.
  • Hamel Husain, Your AI Product Needs Evals — read this before building anything serious.
  • Simon Willison's LLM weblog — the best running commentary on the field.
  • ai
  • llms
  • fundamentals
  • transformers
Need this built? I build ai product or mvp projects for clients worldwide. Tell me about yours.