Most of the people running local AI models today learned the vocabulary the same way. They saw a Reddit thread, copy-pasted a one-liner, picked up a word or two, and started using them in conversations they did not feel qualified to be in. The result is a strange situation where the term in your sentence is doing real work but you have not pinned down what it means. It is not that you are bad at it. It is that the field never slowed down long enough to write a glossary. It just kept shipping tools and trusting the rest of us to figure out the rest.
There is a quiz that caught the gap on purpose. The It’s FOSS team published a ten-question test on the terms that show up in every local AI README, model card, and forum thread. Reports from people who took it suggest the same pattern: the score matters less than the list of words you could not define. The quiz format forces a commitment most explainers skip over. You cannot nod and scroll. You have to pick one. Then live with the answer key when it tells you that the thing you called a token is, in fact, not what you thought it was.
I am not going to retake the quiz with you. The list below is the version I wish I had on day one. These are the words you actually need before you touch Ollama, llama.cpp, LM Studio, or anything that talks to a .gguf file. If half of them feel fuzzy, that is the point. Even people who run these models daily swap the words around in their head. Naming each one cleanly is the difference between picking a model on purpose and downloading whatever file looks biggest.
The unit-economy words you bump into before anything else
Three terms control how much you pay, in time and in RAM. They show up in every model card and every README. You cannot dodge them.
- Token: the chunk a model reads and writes, which is usually part of a word, not the whole word. Common short words like “the” come out as a single token. Long rare words like “antidisestablishmentarianism” split into several. The pricing line “$3 per million tokens” is charging per chunk, not per word.
- Context window: the amount of text a model can hold in working memory at once. Bigger windows let you paste longer documents in. The trade is RAM, latency, and quality loss past a certain length. You will see 4096, 8192, 32768, and 128000 as the common sizes.
- Quantization: a way of shrinking a model by storing the same weights with fewer bits. Full precision is often 16-bit. Quantized versions drop to 8-bit, 4-bit, and occasionally lower. The model gets smaller and faster and pays a small quality hit. When the filename says
Q4_K_M, that is the quantization label.
If those three feel solid, you can read a model card without flinching. If they do not, that is what the quiz is designed to surface. Drop the embarrassment. Saying “I do not actually know what a token is” out loud once is cheaper than six months of squinting at pricing pages.
The format-and-tooling words that decide what you download
The next layer is words that describe the box the model lives in and the harness that runs it. Skip them and you keep downloading wrong files and blaming the laptop.
- GGUF: the file format most local AI tools use for quantized models. If you have ever pulled a model from Hugging Face, the file usually ends in
.gguf. The mental model is shipping container. The contents differ; the format does not. - Harness: the wrapper that runs the model and gives you an interface. Ollama is a harness. LM Studio is a harness. llama.cpp sits underneath, and the tools built on top of it count too. The distinction matters because most “issues” people blame on the model are actually harness behavior.
- Prompt template: the specific wrapper a model expects around your question. Different model families want different wrappers. Get it wrong and the model gets noticeably worse for no obvious reason. That is why tools like Ollama hide this from you. They are doing the version of work the rest of us would forget to do.
- Open weight: a model whose trained numbers (the “weights”) get published for download. Important: open weight is not the same thing as open source. Open weight lets you run it locally, fine-tune it, and ship with it. Open source adds the training pipeline and the data.
Notice the pattern. Every word in this section is something you wish the tool authors had told you about on day one. Most of them did not. They assumed the README was enough.
The work-product words you trip on once you start building
Once you move from downloading models to actually building things, three more terms start showing up in tutorials, blog posts, and the source code of every retrieval-augmented chat app.
- Embedding: turning text into a list of numbers that captures meaning, so a vector database can search it for related content. Two phrases with similar meaning end up with similar number lists. That is how RAG (retrieval-augmented generation, or “look up the right snippet before answering”) finds what to feed back to the model.
- Inference: the act of running the model to produce output. “Inference speed” is how fast it generates tokens. “Local inference” specifically means doing this on your own hardware, which is the only reason you are reading this article in the first place.
- Fine-tuning: taking a base model and training it further on your own data so it gets better at a narrow task. It costs real time and real compute, and most of the time it is the wrong first move. Try the prompt first. Try the harness setting second. Reach for fine-tuning when neither gets you where you need to go.
A useful test for this section: take any tutorial you have been avoiding because of the words in it. Replace “embedding” with “numbers that capture meaning” and “inference” with “running the model” and “fine-tuning” with “extra training on your own data.” If the tutorial is now readable, the writer was showing off. If it is still confusing, the writer was correct and you actually needed the term.
Trade-offs
There is no honest way to write this without naming what learning the vocab costs you.
- The terms are not stable. A label that meant one thing six months ago can mean something else today. The field renames things whenever a marketing team gets involved.
- Forcing yourself to define terms in your own words is slower than nodding along in a chat. It is the only version that actually sticks.
- Knowing the words makes the docs readable but does not make the tools work. You still have to install something, point it at a model, and watch it do the wrong thing three times before it does the right thing.
- Most of the people who use these terms confidently are not smarter than you. They have just been wrong out loud more often.
Where this leaves you
The fastest path through the noise is the same path the quiz hints at. Pick the words you cannot define off the top of your head. There are likely three or four. For each one, write the definition in your own words, not the wiki version. Rewording is the move. It is what turns a term you have seen from a term you can use.
The list of words you cannot define is also the list you bring to the next conversation. Use them. The local AI ecosystem has a documentation problem, and the unofficial fix is people who actually know what they are saying. Saying “I am not sure yet” out loud is more useful than pretending. The vocabulary is short, the tools will keep moving, and the people who bother to learn it now save the people coming after them a lot of time.
Source for grounding: the It’s FOSS local AI jargon quiz and the article published alongside it (linked at the bottom of the original). All ten terms in the explainer section trace to the source list. No first-person use, playthrough history, or unverified mechanic is added.