On this page4 sections
Infermeld lets a Linux user put an existing AMD card and NVIDIA card to work on the same local language model. It launches a separately built llama.cpp server, makes the selected devices and layer placement explicit, and adds preparation checks and a thermal watchdog. Its experimental v0.1.0 release supplies the wrapper’s source, not an inference engine or model weights.
The project appeared in a LocalLLaMA maintainer discussion on October 5. The release record dates publication to October 5 at 02:32 UTC, which was still October 4 in New York. This article examines that release; it is not a new October 9 announcement.
For someone operating a local investigation assistant, the attraction is straightforward: a model that cannot fit on either card individually might become runnable using both. Whether that configuration completes the required investigation quickly enough is a separate question. Infermeld’s recorded experiments help explain where those two questions diverge.
Two cards retain their own memory
A GGUF file packages model weights and metadata for compatible inference engines. In the setup described by Infermeld, one llama.cpp process loads a model and places groups of its layers across two explicitly selected backend devices. AMD uses Vulkan, while NVIDIA uses CUDA. The model’s computation crosses the device boundary as it progresses through those layers.
The pinned llama-server flag reference describes layer splitting as placing layers and their key/value caches across GPUs. Those caches retain attention state for the request. Each card still has its own allocation constraints, and runtime buffers consume memory alongside weights and caches. Adding the capacity printed on the cards does not produce a single interchangeable memory pool.
This makes the placement ratio consequential. Infermeld requires the user to provide two proportions rather than selecting an automatically optimized split. A 3:2 setting is a requested division of offloaded model placement, not a promise that computation or memory consumption will balance perfectly. Different layer contents, cache requirements and devices can change what actually fits.
Consider a hypothetical local assistant that reads a sanitized deployment log and drafts a diagnosis. Its operator has a 16 GB AMD card and a 10 GB NVIDIA card. Choosing 3:2 from those capacities gives the operator a starting hypothesis, but it does not establish that the selected model, context reservation and runtime buffers will fit on each device. The server’s allocation result supplies the next evidence.

A successful load answers one question
The project’s capacity record contains a useful concrete example. On its recorded RX 6900 XT and RTX 3080 pair, the dense Qwen3.8-27B Q4 artifact exited during loading with a 131,072-token context reservation and a 3:2 split. A separate 2:1 attempt at that reservation loaded and passed four short checks.
That result demonstrates placement sensitivity for the exact recorded configuration. It does not establish that 2:1 is a universal best ratio or that the server successfully read a 131,072-token prompt. A reservation allocates capacity for context; a short prompt inside that reservation exercises far less input. The evidence guide explicitly separates allocation and short serving from filled-context ingestion, retrieval quality and sustained throughput.
There is also a separate launcher acceptance record for a pinned Qwen3.6-35B-A3B Q4 artifact at an 8,192-token reservation. The project reports four short checks both with speculative decoding disabled and with its explicit MTP4 setting. MTP uses model-supported draft predictions to propose additional tokens; enabling it requires compatible model data and additional memory. Neither those short checks nor the dense allocation retry qualify another hardware pair.
For the deployment-log assistant, loading gets the operator to a runnable candidate. A representative log packet and a diagnosis with checkable references are still needed to learn whether that candidate is useful. A large reserved context cannot stand in for that exercise.
Inspect the physical devices before the launch
Infermeld’s build instructions pin a CUDA-plus-Vulkan llama.cpp revision. The wrapper needs Linux, Python 3.11 or newer, a separately compiled compatible server and a verified local model. It does not install drivers, fetch weights or assemble the engine for the user.
Device identity is a particularly practical wrinkle. A Vulkan backend can expose an NVIDIA card as well as an AMD card. Consequently, Vulkan0 and CUDA0 being different identifiers does not prove that they refer to two different physical cards. Read the inventory and map the backend names to the intended devices before acknowledging them.
The documented inspection sequence is:
python3 infermeld.py devices --server ./llama.cpp/build/bin/llama-server
python3 infermeld.py serve \
--server ./llama.cpp/build/bin/llama-server \
--model ./models/model.gguf \
--devices Vulkan0,CUDA0 --split 3,2 --confirm-devices \
--ctx-size 8192 --ubatch 32 --dry-run
These paths and identifiers are examples, not a tested recipe for the reader’s machine. The launcher source makes --dry-run emit the proposed argument array without starting inference. The confirmation flag records the user’s acknowledgment; it does not perform physical identification. The preflight tool separately checks required flags, runtime libraries and device inventory. Verify the complete model checksum independently because the launcher’s GGUF magic check does not validate the entire file.
The thermal watchdog reads an AMD junction sensor and stops the server process group that this invocation owns when its guard trips. Its default threshold is 85°C, sampled every 0.1 seconds. Sampling and shutdown take time, so that is a software trip condition rather than a hard temperature ceiling. A missing or unreadable sensor prevents startup or stops the owned server. Adequate cooling remains necessary.
Choose the split using the work that must finish
Our Linux tuning approach compares a bottleneck hypothesis under the same offered workload and keeps latency, completed work and rejections together. Applied here, that means comparing feasible splits using the same engine, model artifact, prompt set, context reservation, output cap and request arrival pattern. Otherwise a change in the test can masquerade as a better placement ratio.
For the hypothetical assistant, use the same sanitized log packets and diagnosis criteria for each candidate. Keep time to first returned text alongside full response latency, accepted diagnoses, failures and guard stops. A faster first fragment is useful during an interactive investigation, but it does not compensate for a diagnosis that never finishes or loses the relevant log evidence. The streaming latency guide explains how to observe those intervals without confusing network chunks with model tokens.
If the required model already runs well on one card, a mixed-device setup needs a workload-specific benefit to justify its added configuration. If it cannot fit on one card, compare feasible mixed-device candidates on their own merits rather than pretending the impossible configuration supplied a performance baseline. Operators who already have a satisfactory llama.cpp launch may also need little of the wrapper.
Infermeld’s useful contribution is making this experiment inspectable: the device selection, launch arguments, retained failures and owned shutdown remain visible. Begin with the model and log packet the assistant actually needs to handle, then let fit and completed-work evidence choose the configuration. That is how two spare cards become a defensible serving choice instead of an arithmetic claim about their combined memory.
Source context
Source links appear within this article. Read them alongside the author’s analysis and evaluate the guidance against your environment.
Report an error or outdated detail