On this page3 sections
A smaller model download is appealing when you want to run AI on your own computer. It may make a model easier to fit alongside the memory needed for your questions and answers. Before downloading it, though, there is a practical detail to check: the file must match the software that reads it and the hardware that runs it. A familiar model name, or even a familiar compression label, does not guarantee that match.
That is the useful starting point for ninfer-ext’s September 29 announcement of two EXL3 builds of Qwen3.8-27B. These are alternate compressed versions of the model, intended for a specific setup: ninfer-ext running on 64-bit Linux with an RTX 5090. They will not load in exllamav3, vLLM or llama.cpp. For someone with the supported machine, the interesting question is how much space the smaller version saves and what, if anything, it gives up in speed. The project’s measurements reveal a useful tradeoff.
What the EXL3 label tells you
A model’s weights are the numbers used in its calculations. Quantization represents those numbers with fewer bits, reducing the amount of data the engine must store and read while introducing an approximation. The engine needs to understand that representation. ninfer-ext, a fork of Neroued’s NInfer C++/CUDA engine, has its own EXL3 implementation: a compact, trellis-based encoding with matching GPU kernels, the routines that decode the weights during computation. Its quantizer starts from full-precision tensors rather than importing an exllamav3 checkpoint.
The resulting download is a .ninfer artifact, meaning a model file packaged for this loader. The model card is explicit that these artifacts require ninfer-ext; Transformers, vLLM, llama.cpp and exllamav3 cannot load them. This is the point that caused confusion in the Reddit comments. The repository is Apache-2.0 licensed, so an engine-specific format should not be confused with closed-source software. Licensing and compatibility answer different questions: whether you can use or modify the software, and whether your chosen engine knows how to run the file.
There is a third layer that can make the setup look more interchangeable than it is. Once running, ninfer-ext provides OpenAI- and Anthropic-compatible interfaces for applications to send requests. A documentation assistant may therefore be able to keep much of its existing client code while changing the server it contacts. That convenience does not extend to the model file underneath. The client speaks to the server; the server still needs the right loader, weight encoding and GPU. Separating those layers helps explain why a familiar API can work even though a familiar model download cannot.
Read diagram description
First match the .ninfer model file to ninfer-ext. Check its current Linux and RTX 5090 requirements. Account for model weights, the conversation cache and temporary working memory. Then compare acceptable output quality, required context and concurrency, and latency under the same workload. A smaller file does not guarantee faster work.
Where the smaller file helps, and where it costs time
Consider a hypothetical team using a local documentation assistant on one Linux workstation with an RTX 5090. Most questions are short, such as finding a setting in a reference document. Occasionally, an engineer supplies a long configuration bundle and asks the assistant to explain a mismatch. The team is considering a smaller model artifact because it wants more room for those longer requests, while keeping everyday answers comfortably responsive.
There are two published choices: 4.0 bits per weight, or bpw, at 15.68 GiB, and 3.5 bpw at 14.24 GiB. The bpw labels describe the main weight encoding; some components use other rates. Both include text, multi-token prediction and vision support. Neither the label nor the file size tells the team its total GPU-memory requirement. The running engine also needs memory for execution and the KV cache, which retains attention state from the request so the model can reuse it while generating subsequent tokens.
Saving weight space may therefore be useful, but it is only part of the calculation. More compact weights can also require more work to decode. In this implementation, the model card attributes the 3.5-bpw build’s lower speed to that extra work at half-bit rates. The smaller file reduces data movement, yet the decoding cost more than offsets that saving in the reported tests. Here are two of the project’s measured comparisons:
| Published artifact | File size | Plain decode | MTP, three draft tokens |
|---|---|---|---|
| EXL3 4.0 bpw | 15.68 GiB | 75 tokens/s | 137 tokens/s |
| EXL3 3.5 bpw | 14.24 GiB | 64 tokens/s | 130 tokens/s |
Decode is the stage that produces the answer token by token. MTP stands for multi-token prediction, used here for speculative decoding. These are project-reported single-request spot measurements on one RTX 5090 with CUDA 13.3, greedy generation, 64–256 output tokens and a prefill chunk of 1,024 tokens. Prefill is the initial processing of the prompt. The figures offer a reason to compare the two artifacts, but they do not tell the documentation team how several simultaneous questions will behave or whether the answers will meet its needs.
For the team’s short reference questions, the larger artifact may already fit easily and respond faster. The smaller one becomes more interesting if it lets the service handle a longer request that otherwise cannot fit. That is a different benefit from maximizing tokens per second, and it deserves a different test. The team needs to find out whether the saved space changes a real request outcome, then decide whether any additional waiting is acceptable.
Try it on the work you actually do
The first step is to establish a working combination. The repository requires an RTX 5090 with sm_120a support and a suitable CUDA toolkit; CUDA 13.3 is listed as validated. Its README gives the Linux build dependencies and instructions. Other GPUs and community ports are outside the project’s stated support, even when a discussion reports success with them. After building ninfer-ext and downloading the matching 4.0-bpw artifact, the model card gives this starting command:
build/apps/ninfer-serve qwen3_8_27b_exl3_4bpw.ninfer \
--port 8099 --spec mtp --draft-tokens 3 --fixed-draft
This is a documented launch example, not an AIOpsSRE runtime test. Record the engine revision, artifact and settings so the comparison can be repeated. Successfully loading the model establishes that those pieces work together. To decide whether to replace an existing configuration, the documentation team still needs the short questions and long bundles that motivated the change.
Write the intended improvement before running that comparison: enough memory for the required longer requests, without an unacceptable delay in ordinary answers. Keep the same questions, input lengths, output limits and arrival pattern for each candidate. Preserve the cache settings and speculative-decoding mode in the results so a configuration change does not get mistaken for a format improvement. This follows the method in our Linux performance guide, where the measured bottleneck determines the experiment.
The useful record includes time to the first token, total completion latency, completed requests and requests rejected because of resource limits. Checking those together matters: a service can look faster simply because it refuses more of the difficult work. The team should also compare answers against the same known references. An answer that arrives quickly but misses the configuration exception the engineer needed has not completed the job. Our metrics guide explains why the request population and measurement window belong beside the numbers.
If both artifacts handle the required inputs, the smaller download may offer this team little practical advantage. If its reduced footprint makes a previously failing request feasible, accepting slower generation could be worthwhile. Keep the configuration that serves those requests well, with enough information to reconsider it when the workload changes. That is what the new files make possible: a specific choice about a supported local setup, grounded in the work the assistant needs to do.
Sources & context
Sources linked in this article. Read alongside the author’s analysis; a citation does not independently verify a publisher’s claims.
- ninfer-ext’s September 29 announcementwww.reddit.com
- ninfer-extgithub.com
- model cardhuggingface.co
