Make local models usable before making them bigger.
Notes from trying Ollama, quantization, VRAM, and context—moving local AI from a downloaded model toward a usable work layer.
RESEARCH NOTE01
Ollama was an entry, not the only source.
Once a model ran, I separated where it came from from how it ran.
I do not always get models through Ollama. I often choose a model or GGUF file directly from Hugging Face, then decide which runtime fits it. Ollama is useful because downloading, managing, and exposing a local API are easy entry points.
Once I cared about quantization formats, context, GPU memory, and how inference actually runs, I moved deeper into llama.cpp and GGUF. The model source, file format, runtime, and agent are separate decisions; getting a model to start is not the same as finishing work.
My main setup is a Windows desktop with an RTX 3080 Ti, 64GB of RAM, and an SSD. With models such as Qwen3.8-27B, I look at weights, KV cache, context, and the headroom needed by other applications. An SSD helps reads and writes; it does not add VRAM.
- 01
Prepare
Get the model
- 02
Run
Files and tools
- 03
Check
Can work continue?
02
Fitting the file is not enough.
Budget weights, KV cache and runtime space.
I used to focus on file size. A compressed model close to VRAM capacity seemed usable, but execution also needs KV cache, workspace and buffers. Windows and other apps need room too.
Agent context accumulates files, diffs, tool descriptions and terminal output. With longer context and more sessions, I experienced rising memory use, desktop stalls, driver resets and Windows freezes. These happened in my environment; they are not universal outcomes.
I now leave headroom instead of opening the maximum context immediately. Confirming one stable task before adding context or concurrency is more practical for me.
| STEP | Weights | Context | Headroom |
|---|---|---|---|
| Model data | Included | Not included | Not included |
| KV cache budget | Not included | Included | Not included |
| Runtime and other apps | Not included | Not included | Included |
03
Bring Q4, Q8 and IQ back to the same task.
Size and usability need separate checks.
Q4_K_M, Q8 and IQ variants kept appearing in my GGUF research. I initially ranked them by bit count, then started checking format, runtime support and remaining memory.
Smaller IQ3 variants were appealing because larger models seemed closer to my hardware. But an agent still carries long context after weights shrink. A smaller download does not establish long-running stability.
I use llama.cpp / GGUF as a research baseline. This is not a completed quantization ranking. Next I want a fixed task and evidence that edits and checks finish, rather than just fluent answers.
- Size
- Weights and headroom
- Support
- Runtime compatibility
- Finish
- Complete the same task
04
More quantization names do not make a better setup.
Separate method, numeric format and runtime.
Researching AWQ, GPTQ and QAT showed me that quantization is not explained by a 4-bit label. The methods handle error differently, and the target runtime still has to support the resulting format.
I also compare original weights and GGUF files directly on Hugging Face. Unsloth’s Qwen3.8-27B-GGUF is a concrete route to a pre-quantized Qwen3.8-27B file, so I can think about file size, remaining memory, and execution instead of treating quantization as only a theory.
I do not treat Unsloth as a harness or claim that I trained and ranked the model. It is one way to obtain and compare a quantized file; after downloading it, an inference environment such as llama.cpp or vLLM still has to run it against a real task.
- Method
- AWQ / GPTQ / QAT
- Format
- FP8 / NVFP4
- Execution
- Check runtime and GPU
05
Watch the wait before the answer too.
Separate loading, prefill, TTFT and decode.
Prefill processes the prompt and context; decode generates output. TTFT is the wait for the first token. Fast output does not remove a long wait before it starts, which can interrupt my work.
After reading files or running tools, an agent brings new content into the next step. My next comparison should record loading, input processing, first output and generation separately, including longer context, rather than compressing everything into tokens per second.
- 01
Load
Prepare model
- 02
Prefill
Process input
- 03
Decode
Record TTFT and generation
06
NInfer and Strata widened the questions.
Keep research directions separate from measurements.
NInfer drew my attention to matching particular models, hardware and runtimes. Strata made me consider how a whole PC could share work when a model does not fit the GPU. My interest moved beyond choosing a model.
These are research directions. Project demonstrations are not results from my Windows desktop. Deployment tools also need to fit the hardware, so I am comparing Ollama, llama.cpp / GGUF, LM Studio, and vLLM around startup, APIs, context, tool calls, and stability. These are research configurations, not one finished deployment.
NInfer
Study focused configurations
Strata
Study whole-PC resource use
07
Next, compare the same task through completion.
Move from speaking speed to the result.
I want to hold the repository, task and tests constant, then vary model, quantization, runtime or harness. Changing one main condition at a time should make differences easier to trace. This is my next testing plan.
Alongside time and memory, I want correct tool calls, understandable edits, tests actually run and a record of human intervention. A usable final result is closer to my daily needs than the amount of generated text.
- 01
Fix
Repository, task, tests
- 02
Observe
Time, resources, tools
- 03
Accept
Is the result usable?
08
Separate the model, runtime, service, and harness.
Draw the work boundary instead of ranking names.
I now separate local AI into four roles: the model file, the runtime that executes it, an API or service when other software needs to call it, and the harness that reads a project, uses tools, manages context and permissions, and returns work for acceptance. These are roles in a workflow, not products on one ranking.
My model sources are mixed. I often go directly to Hugging Face for original weights or GGUF; Ollama is an easier entry for downloading, managing, and exposing a local API; llama.cpp is where I study GGUF, quantization, and GPU execution more closely; and vLLM is an inference engine oriented toward serving and throughput. llama.cpp can expose a server API itself, so service is a deployment role rather than a mandatory extra product.
OpenCode and Codex are the tools I actually rely on for daily work. Pi, DeepSeek Harness, and MiniMax Code are exploratory tools I use to understand design differences. I look at how they connect models, context, and tools, but I do not present a trial as daily experience.
- Model
- The weights
- Runtime
- llama.cpp, GGUF, inference
- Service
- Ollama, vLLM, API
- Harness
- Context, tools, acceptance