↓Skip to main content
  1. /classes/
  2. Classes, Fall 2026/
  3. CS 2010 Fall 2026: Course Site/

cs2010 Notes: 09-30 Large Language Models

·649 words·4 mins·

LLM Training
#

Last time we ended at LLM training:

  • Pick an LLM archetecture and size (e.g. dense 9B, delta-net linear attention)
  • Get enough training data, at least 20 tokens per weight (e.g. 200 GB of text, or 10x the size of Wikipedia)
  • Feed in token 1, it predicts token 2. It’s wrong, but we can compute whether each weight should have been higher or lower go get it right. Tweak each weight slightly.
  • Repeat for all the tokens.

Training phases:

  • Initially a “base” model is trained, that just predicts continuation of sequential text.
  • For a chatbot, you want it to predict text within a turn-based chat sequence and thus respond to a user’s prompts. This is done in a second training/fine tuning round that produces an “instruct” variant.
  • After that, variants can be created through additional training. For example, you could take an existing model and train it to talk like a pirate.

Conversation Protocol
#

Nest demo, chat, Qwen 3.8, thinking off.

  • System prompt.
  • User / assistant alternation.
  • Show API messages, talk about JSON.

Nest demo, Coding, Qwen 3.8, thinking on.

  • Thinking tokens.
  • Tool calls.
  • Tool specifications.
  • API details.

What can we do with tools?
#

  • We saw this in lab on Monday
  • Write HTML, markdown, convert between text formats, etc.
  • Use Pandoc to convert markdown to PDF.
  • Write simple computer programs.

Running LLMs
#

LLMs run on computers. You need a program (an inference engine) that will load the model and run it.

How big an LLM can you run?

  • LLMs have some number of weights (parameters) (e.g. 4B, 397B-A17B)
  • A full size weight is a 16 bit floating point number, so one weight “naturally” takes 2 bytes to store.
  • Weights can be “quantized”, or stored with less detail, without messing things up too much. Models are typically run in quantized mode with 8 bits per weight for good quality or 4 bits per weight for okay quality.
  • At 8 bpw, 1 weight = 1 byte, so a 4B model takes ~4GB.
  • At 4 bpw, 2 weights = 1 byte, so a 4B model takes ~2GB.

A recent technique is quantization-aware training, where a late training phase runs at q8 or q4 specifically to improve quality when running at that quant.

To run fast, you want weights to fit in video memory.

You also need to fit something else in video memory: KV cache. That’s the already-processed active conversation, so every new chat turn doesn’t need to reprocess the whole thing from the beginning.

How big are these models? What do they run on?
#

  • Current top is 5+T models like Claude Fable, OpenAI Astra
    • A full rack of servers (draw 3ft by 2ft by 8ft high)
  • Below that is 1-3T models like Kimi K3
    • One big server (draw 3ft by 2ft by 2ft high)
  • Below that is 200-800B “Flash” models like Deepseek V4.1 Flash
    • One medium server (draw 3ft by 2ft by 2ft high)
    • One maxed out workstation
  • Below that is small general purpose models: 25-150B weights, like Qwen 3.8 27B
    • These can run on a small server or nice workstation
    • Small ones can run on absolute maxed out gaming PCs (e.g. RTX 5090)
  • Below that is very small models: 0.5-15B weights
    • Run even on old gaming PCs
    • Like we ran in lab yesterday
    • Potentially useful, but mostly you need to tell them exactly what you want to get anything useful out of them

Exploring Performance
#

  • Prompt processing
    • Parallel computation, can run all the prompt tokens through the layers of the model at once
    • How much compute power do you have?
    • Really want 1000+ for usability, especially on larger prompts
  • Token generation
    • Need to run one token through the whole model
    • How much memory bandwidth do you have?
    • Not too bad at 20+.

GPU vs. Unified Memory vs. CPU

Do a CPU vs. GPU llama-bench test on Kraken. Talk about multi-channel memory and stuff.

Nat Tuck
Author
Nat Tuck