Qwen 3.8 Splash, served locally
A 27B model on the Mac itself, with no per-token bill.
What it is
This build runs a private AI assistant on a 64 GB Apple M5 Pro Mac. The second Hermes Agent profile uses a 4-bit build of Qwen3.8-27B, served by the Inco Splash 1.0 engine, listening only on the Mac itself through an OpenAI-compatible API. It is the profile's default model for Telegram chat, its sub-agents and 16 types of background task, though from 2026-09-24 to 10-02 most chat ran on a second local runtime. It replaced an Ollama build of the same model family on 2026-09-19 after a benchmark showed about double the speed.
By the numbers
What it means for your team
A small office with a 64 GB Mac can run a private assistant that keeps data on the machine and has no per-token bill. It needs monitoring, and it should be tested against hosted models before it takes on work where quality matters.
Status
Quality was checked on two real tasks. In Maillard, plans drafted on Splash passed the 12 food-safety cases the fixed rules enforce, but the full eval failed on side-dish coverage. On 24 TrimIndex buyer briefs, Splash passed 8 where Claude Sonnet had passed 23. Memory on a 64 GB Mac is a real limit: a second local model kept loaded beside Splash left too little room for other work, so it was switched off on 2026-10-08. The model and engine are third-party; Nix Inc built the serving setup, the watchdog, the benchmarks and the routing.
Serving
Splash runs the 4-bit Qwen3.8-27B model with a draft model of about 1.3 GB for speculative decoding, a 64K-token context and a 40 GiB memory cap. A wrapper script starts the server directly because the stock launcher's port was already in use.
Running as a service
A launchd job starts the server at login and restarts it if it stops. After a 1h44m outage on 2026-09-21, a watchdog was added that sends a one-token test request every 120 seconds and restarts the engine only on the failure seen in that outage. It has fired once, on 2026-09-23.
Measured settings
The same benchmark, one warm-up and three 512-token runs, ran on the old and new runtimes, then with four simultaneous requests. A reasoning probe at three effort levels chose medium for chat, correct in 19.1 s against 40.3 s at the highest level; the highest level is kept for delegated sub-agent work.
Safe config changes
Scripts rewrite the agent config by making a timestamped backup, writing a temp file and renaming it in one step. The previous model was kept, so rolling back is a single step.
Hardening
After memory stalls on 2026-09-20 and 09-21, 222 copied skills were disabled, leaving 50, and a fallback model that had loaded beside Splash was removed. A cold call fell from 22,117 to 16,663 tokens and from 55 s to 39 s.