Practical inference guide
Local AI: fit the model, then measure the workflow
A large memory specification can make an interesting development machine. It does not tell you whether a particular assistant feels responsive. We have not run original RTX Spark inference tests; this guide explains how to evaluate claims and collect results without inventing tokens/second numbers.
What the platform claim does—and does not—mean
NVIDIA advertises up to 128 GB of unified memory, native CUDA support and up to one petaflop of FP4 AI performance. Its specification table also distinguishes laptop configurations with different maximum memory capacities.[3] These are vendor specifications, not measured application throughput. Peak FP4 arithmetic is not actual tokens/second: the model format, kernels, memory behavior and workload still need to be established.
Treat the widely repeated 120B-parameter local-model statement as a vendor capability claim, not a measured speed or a guarantee for your chosen model.[5] Its exact model, quantization, context and software conditions have not been established in this guide; not all SKUs have the memory capacity implied by that claim. Do not purchase a lower-memory configuration on the assumption that a maximum-capacity demonstration applies unchanged. Even a model that loads can be too slow for an interactive task.
1. Define the model and memory budget
Write down the exact model repository, revision, architecture, parameter count, file format and quantization. A model family name is not enough: different weight formats can change memory requirements, supported kernels and output quality. Evaluate the quantized model against examples from your intended task, not just whether it produces text.
Unified RAM is a shared budget, not a promise that every installed byte can hold weights. Reserve space for the OS, running apps, backend allocations, work buffers and KV cache. The KV cache holds attention state and its size depends on model architecture, context, cache format and concurrent requests. Test the context you actually need, including system prompts, retrieved documents and generated output. A short-prompt success is not proof that a long-document workflow fits.
Watch peak memory during loading and during generation. If you approach the memory limit, reduce context or concurrency, try an appropriate smaller model or supported quantization, and retest quality. Do not assume swapping to storage preserves acceptable interactivity. Keep disk capacity for downloaded weights, caches and additional model revisions in the buying checklist.
2. Verify the entire backend, not just CUDA branding
Choose an inference backend with documented support for your OS, Arm64 CPU architecture, GPU and model format. Record its version or commit, driver version, runtime and compiled dependencies. The native CUDA platform claim does not automatically supply a working native build of every existing x64 application or extension. See the compatibility guide before transferring a desktop-PC environment.
Inspect backend logs to confirm which device performs the work and whether layers or operations fall back to the CPU. Record GPU offload, batch settings and cache precision. A Vulkan run, a CPU-only run and a CUDA run are separate setups, not interchangeable results. Prototype driver failures are also not a retail verdict; our hands-on hub keeps sample limitations attached to community reports.
3. Measure prefill, decode and waiting time separately
The llama.cpp benchmark documentation separates prompt-processing and text-generation tests and records backend, model, build and test parameters in its output.[4] Use the documentation for your installed version; a command from another release is not a reproducibility guarantee.
- Prefill: processing the input prompt. Record input token count and prompt-processing tokens/second; do not report this as output generation speed.
- Decode: generating new tokens. Record output tokens/second, output length and the context already in use. Distinguish one user's response from aggregate batched throughput.
- Time-to-first-token: the wait from submitting a request to receiving its first generated token. Measure at the application boundary and state whether loading, queueing and prompt processing are included.
- Cold versus warm: record initial model load separately from repeated requests. Repeat the same workload and retain spread and failures, not only the fastest run.
Record charger state, selected power profile, configured power limit where available, sustained power draw and thermal conditions. Battery and plugged-in sessions must be labeled separately. Repeat a longer representative task to find throttling or instability; a brief synthetic run can miss both. Include task quality and successful completion alongside speed.
A result card you can actually compare
Capture device and memory SKU; OS build; driver; backend build; model revision and quantization; context and cache format; input/output lengths; batch and concurrency; offload settings; prefill and decode rates; time-to-first-token; peak memory; power mode and measured power; repetitions; errors; and the original log. Leave missing fields unknown instead of filling them with platform specifications.
| Measurement | RTX Spark evidence in this guide |
|---|---|
| Prefill / decode tokens/second | Pending reproducible retail tests |
| Time-to-first-token | Pending disclosed application measurements |
| Usable model capacity | Depends on SKU, quantization, context and memory reserve |
For local agents, add a safety check: confirm whether the application actually stays offline, where logs and documents are stored, and which tool permissions it receives. Local inference does not make a network-connected tool private by default. Start with a small, constrained workflow and expand only after reliability and access controls are understood.
Sources
[3] NVIDIA: platform specifications and native CUDA claim
[4] llama.cpp: llama-bench documentation and output fields
[5] NVIDIA announcement: local model capacity claim, not a measured inference result
Check the exact device and memory configuration →