This is a pure C/Metal inference engine for running 397B parameter MoE models on Apple Silicon.
metal_infer/
infer.m # Main inference engine
shaders.metal # Metal compute kernels
chat.m # Interactive chat TUI
main.m # MoE-only benchmark
tokenizer.h # Single-header C BPE tokenizer
linenoise.c/h # Line editing library
modelmgr/ # Model management core (registry, manifests, recipes,
# artifact hashing/verify, offload, migration, TUIs)
# — see docs/MODEL_FRAMEWORK.md
steps/ # Shared step library (download, tokenizer, extract,
# repack, compile_native, materialize)
tui/ # Onboarding, config wizard, manage flows
assets/models/ # Per-model manifests; assets/model_configs.json is
# GENERATED from them (`make registry`)
tests/
test_flashchat_cli.sh # CLI regression test
test_modelmgr_cli.sh # Model management integration test
test_api_smoke.sh # HTTP API smoke test
python/ # modelmgr unit tests (make py-tests)
root/
Makefile # Project build/test surface
assets/api_perf_log.tsv # Durable API perf trend log (make bench-api)
# From the project root
# Build all targets
make
# List available Make targets
make help
# Build specific binaries
make metal_infer # Main benchmark
make infer # Full inference engine
make chat # Interactive chat TUI
make api-smoke # HTTP API smoke test
# Run benchmarks
make run # Single expert forward pass
make verify # Metal vs CPU reference verify
make bench # Benchmark single expert (10 iterations)
make moe # Full MoE forward (K experts, single layer)
make moebench # Benchmark MoE (10 iterations)
make full # Full model forward (K=4)
make fullbench # Full forward benchmark (3 iterations)
# Inference engine (full model)
./metal_infer/infer --prompt "Hello" --tokens 100
# With timing breakdown
./metal_infer/infer --prompt "Hello" --tokens 20 --timing
# Framework regression tests
make cli-smoke # tests/test_flashchat_cli.sh
make manage-smoke # tests/test_modelmgr_cli.sh
make py-tests # tests/python/ (modelmgr unit tests)
make registry-check # assets/model_configs.json matches assets/models/*.json
make tool-template-smoke
make api-smoke
make test # everything
# Performance regression benchmark (server-level, per model) + trend report
make bench-api
make bench-reportmake bench-api is the canonical perf-regression command. It derives its matrix from
the model registry (assets/model_configs.json) — every installed model not opted out
with "benchmark": false is benchmarked in its as-shipped default config, against the
real HTTP server, with one uniform spec (fixed prompts, temp 0, warmup + repeats),
recording prefill (TTFT) and decode (tok/s) separately to assets/api_perf_log.tsv.
make bench-report compares the latest rows to prior commits (keyed on hw_model) and
flags regressions.
Do not selectively revert or omit assets/api_perf_log.tsv rows produced by normal
validation commands such as make api-smoke, make bench-api, or related test
surfaces. The log is a watchdog/sentinel artifact: its value is time-inverted because
future regressions may need historical rows that did not look important when they were
captured. If normal testing generated rows, preserve them unless the user explicitly
asks to discard them or the rows are known corrupt/non-test output. When committing a
verified unit of work, include assets/api_perf_log.tsv by default if it changed; do
not treat it as incidental dirty-worktree noise or exclude it from the commit just
because the code change lives elsewhere.
This design exists because the perf log silently went stale for the entire MTP + dense arc: the suite followed a single configured model, so new models/quants were never measured. Coverage is now structural, not manual — but two things still require a human:
- Adding a model/quant variant: put it in the registry (you must anyway, for the TUI)
and it is benchmarked automatically. Only set
"benchmark": falsefor variants that physically cannot run here (e.g. the 397B). Do not add a parallel benchmark list — a second source of truth is exactly what rotted last time. - Adding a perf-affecting axis that is not a model (a new kernel mode, an MTP
depth, a sampling/runtime toggle that ships on by default): it is exercised
automatically only if it's part of the model's default config. If it's a new opt-in
path, add it to the spec in
tests/bench_api.shso every model covers it.
Run make bench-api (or at least --model-id <id>) for any change to the decode/prefill
hot path, kernels, attention, or speculative decoding, and confirm make bench-report
shows no regression before committing. Ad-hoc --mtp-generate-* numbers are for
inner-loop iteration, not regression sign-off.
For dense prefill performance work, investigate Metal 4 MPP/TensorOps matmul2d
for the gpu_dequant_matmulN / gpu_dequant_matmulN_batch path first. Treat
MPSNDArrayQuantizedMatrixMultiplication as a useful affine-quantized comparator
or fallback probe only if the local SDK exposes it, but keep the likely production
direction focused on a low-overhead MPP kernel path, gated by runtime capability
checks and preserving the existing Metal 3 kernels as the fallback on older
toolchains or devices.
Preserve MLX affine quantization semantics: TensorOps experiments must process
64-wide quant groups and apply each group's scale/bias with a per-token group
sum; a single full-K raw uint4 matmul is not equivalent. BF16 activations are the
current production-oriented input candidate, and zero-copy integration still
needs an aligned/repacked 4-bit tensor layout or another validated binding path.
Current SDK probes also show two practical MPP limits: Metal buffer argument
indices top out at 30, so group-wise TensorOps kernels must chunk groups, and
the simple one-MTLTensor-per-group shape is only validated for small output
tiles before TensorOps returns zero results. Do not wire this into production
until a compact resource layout or tensor-pool strategy is validated at full
80-group dense dimensions.
Rule: every debug/profiling feature added to the engine MUST be exposed as a
setting in the interactive config menu — not env-var-only. A flag a user can't
reach without exporting FLASHCHAT_* by hand is effectively undiscoverable.
The menu's advanced section is the home for these toggles:
- Menu UI:
modelmgr/tui/config_wizard.py→_advanced_settings()("Configure advanced options (debug, MTP, cache)?"). Add a(KEY, "label", "default")tuple here.
A new debug key must be wired through the whole config chain, or it won't reach
the engine. Use an existing debug flag (e.g. SERVER_HTTP_LOG) as the template —
grep it to find every site:
lib/config.sh— eight sites:FLASHCHAT_DEFAULT_<KEY>constant, the reset/var declaration (top-level + insideflashchat_load_config), the_flashchat_migrate_configbackfill list, theFLASHCHAT_<KEY>env-override apply, the finalize-default line, and theflashchat_getgettercase. BumpFLASHCHAT_CONFIG_SCHEMA_VERSIONso existing user configs backfill the key.flashchat— thecmd_servelaunch block bridges config → environment (FLASHCHAT_<KEY>="$(flashchat_get <KEY>)" \on theinferexec line); runtime settings are read from the config file/env, not shell args. Also add the key to theconfig --showdump.metal_infer/infer.m— the engine reads the flag viagetenv("FLASHCHAT_<KEY>"). Treat empty string as "unset/default" (an unset config key bridges through as""), not as a literal value.
Reference implementations — both added across exactly these sites, use either as the worked example when adding the next debug feature:
- pread timing profiler —
FLASHCHAT_PREAD_PROFILE/FLASHCHAT_PREAD_PROFILE_CAP(prof_preadwrapper ininfer.m; analyze withtools/pread_profile_analyze.py). - expert pin-cache —
FLASHCHAT_EXPERT_PIN_MAX_GB(cap GiB, 0=off) /FLASHCHAT_EXPERT_PIN_AUTO_FRAC(free-RAM fraction) /FLASHCHAT_EXPERT_PIN_MLOCK. Keeps the hottest whole experts RAM-resident (LFU,expert_pin_*ininfer.m) so low-RAM machines don't thrash the page cache re-streaming experts per token. Budget =min(AUTO_FRAC × free RAM at first use, MAX_GB), fail-soft to pure pread.
- Write self-documenting code with clear naming
- No comments unless required for the complexity (comments are never for historical context)
- Prefer early returns for error conditions
- Keep functions focused and single-purpose
- Configuration policy: shipped model metadata belongs in
assets/model_configs.json; user settings belong in~/.config/flashchat/configand should select models byMODELID; generated per-model runtime artifacts belong in<model>/flashchat/so entire model snapshots can be moved or restored as a unit. Model offload storage is one globalOFFLOAD_DIR, never a per-model config path. - User/app state policy: Flashchat-owned user state belongs under
~/.config/flashchat/, including sessions, prompt history, default server logs, pid files, and optionalsystem.md. - Do not add implicit config fallbacks. Use
--config FILEonly as an explicit override, otherwise use~/.config/flashchat/config, environment overrides, and defaults derived from the bundled model registry. - Do not rely on current working directory for model registry lookup. Shell, Python, and C/Objective-C callers should resolve the registry via
FLASHCHAT_MODEL_CONFIGor the repo-rootassets/model_configs.jsonpath. - Only the registry-backed production expert format is supported in active code. Old low-precision experiments may remain in historical results/paper artifacts, but do not reintroduce alternate runtime paths, setup scripts, config settings, or user-facing UX.
- Model storage management must be registry-aware and confirmation-gated. Whole-model offload/reload operates on HuggingFace cache repo directories, runtime-only restore copies only
<model>/flashchat/, and destructive model actions require typing the exact model ID. - Always ask the user before modifying their system (e.g., installing packages, changing config files, running system commands)
- Setup/dependency installs must start with one explicit consent screen (
Y/x) before any download/install work begins.Xmust cancel cleanly with a clear "User cancelled" style message and instructions to re-runflashchatlater. - When the user gives specific instructions for moving forward, ask if you should update the AGENTS.md file to ensure the instruction is adhered to.
- When testing interactive scripts (especially with piped input), if output is confusing or unreadable, improve the output formatting for clarity before continuing
- Fresh-install/setup testing must use isolated timestamped
HOMEdirectories under the project, such asdebug/fresh-envs/YYYYMMDD-HHMMSS/home. Do not test setup flows by toggling, renaming, or moving the user's real~/.config/flashchator model cache. - Never assume the user has knowledge of the system, commands, or syntax. Always provide instruction that is atomic and self-explanatory.
- Use user-friendly terminology, not technical precision. For example, use "context window" instead of "max tokens" since that's what users expect in modern LLM interfaces, even though the technical term is different.
- When making user-facing changes (especially output/UX), look for similar patterns elsewhere in the project. For example, if you consolidate section headers in one command, check other commands for the same issue and apply the same fix.
- Don't ask the user to perform manual steps that can be automated. If something needs testing, write a script to do it automatically rather than asking the user to do it manually.
- When changing the HTTP/API surface, add or update an automated smoke test whenever feasible. Prefer a lightweight script that exercises the live endpoints the same way frontend harnesses will.
- Tool-calling changes should validate rendering before live agent testing. Use
metal_infer/infer --render-request ... --render-output ...andmake tool-template-smoketo verify native Qwen tool-template rendering and XML parser behavior before comparing nanocode/opencode live logs. - Shell helpers should preserve the caller's working directory unless changing directories is the explicit purpose of the helper. Save and restore
pwdinside build/setup helpers thatcdinternally. - Always verify the environment state before and after making changes. Check that config files, model caches, and other state are preserved or properly restored. Test that changes don't inadvertently delete or corrupt user data.
- Server-side debug visibility matters. Changes to
infer --serveshould preserve persistent logging so background server runs remain debuggable without requiring an interactive launch. - Persistent system prompt caches belong with the model runtime artifacts. Store them under
<model>/flashchat/system_prompt_cache/unless set to an alternate path via the user config. validate model dimensions and prompt hash before reuse, write atomically, and keep cache counts bounded to avoid unnecessary SSD churn. - Running server state must be invalidated when runtime inputs change. If model selection, model registry contents, server-affecting config, or infer source/binary state changes, Flashchat-owned servers should be treated as stale and restarted before reuse.
- When runtime behavior does not match expectations, anchor debugging in captured logs, request dumps, build stamps, and reproducible evidence. Do not default to assuming the user launched an old binary or made a procedural mistake unless the evidence specifically points there.
- Server control commands must verify reality, not assume it. Start/stop helpers should confirm health or actual process exit before reporting success or removing pid/state files.
- Persistent server/runtime toggles should be first-class config options. If a setting is useful beyond one-off debugging, surface it through the config file and
flashchatconfiguration wizard instead of leaving it env-only. - Sampling/runtime generation knobs must stay consistent across config, API request parsing, server runtime signatures, docs, and smoke-test perf logging. Do not add hidden sampler defaults that cannot be inspected or overridden from Flashchat's normal user surfaces.
- Model-specific sampling profiles belong in
assets/model_configs.json. The Flashchat config should select aSAMPLING_PROFILE, withcustomreserved for manually managed sampler/reasoning values. - Multi-token prediction policy: MTP is a model/profile/server-default capability, not just a debugging flag. Treat user config
MTPas a numeric draft budget: empty means use the registry fallback chain,0means force disabled,1means automatic, and2+means that predictor batch size. Model and profile metadata usemtp_default_predictions; profile values override server defaults, server defaults override model defaults, and model defaults should be1unless there is a specific reason otherwise. MTP-capable models can require extracted MTP artifacts independently of the selected runtime budget. - Python interpreter resolution must go through
lib/python.sh. The launcher never calls barepython3for venv creation or package install — alwaysflashchat_python_bin(or its resolved path). On macOS 26,/usr/bin/python3is CLT 3.9 and will shadow any Homebrew Python on PATH;python3 -m venvagainst it produces a 3.9 venv that cannot importmodelmgr/because the codebase uses PEP 604 union syntax unconditionally. Minimum supported interpreter is 3.10; the resolver enforces this and refuses to fall back silently. - Stale-venv recovery is silent. If an existing
metal_infer/.venv/interpreter falls below the supported Python floor,ensure_venvrecreates it with a one-line log message — no user prompt, no consent screen, because nothing on the user's system is being modified (Homebrew Python lives under/opt/homebrew/, not/usr/). The venv is regenerated state, not user data. - When a verified coherent unit of work is complete, ask whether it should be committed before ending the turn. Don’t interrupt active debugging for every small change, but after implementation plus relevant verification, make the commit checkpoint explicit instead of leaving it implicit.
Indentation: 4 spaces (no tabs)
Naming Conventions:
- Global variables:
g_<name>(e.g.,g_timing,g_cache_telemetry_enabled) - Constants:
UPPER_SNAKEorkCamelCase(e.g.,MAX_K,kMetalDevice) - Functions:
snake_case(e.g.,load_expert_weights) - Types:
snake_case_torCamelCase(e.g.,bpe_tokenizer,LayerTimingAccum)
Imports:
// Objective-C (Foundation, Metal)
#import <Foundation/Foundation.h>
#import <Metal/Metal.h>
// C standard library
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <stdint.h>Error Handling:
- Return error codes (0 for success, -1 for failure)
- Check return values immediately with early returns
- Log errors to stderr with descriptive messages
- Use
NSError **for Objective-C methods
File Structure:
// ============================================================================
// Section divider comment
// ============================================================================
static helper_function() { ... } // file-scope functions are static
// Public function
int public_function(int arg) {
if (arg < 0) {
return -1;
}
// ... implementation
return 0;
}Indentation: 4 spaces
Header Format:
/*
* shaders.metal — Brief description
*
* Detailed description of operations.
*
* Quantization format: ...
*/
#include <metal_stdlib>
using namespace metal;
// ============================================================================
// Kernel: kernel_name
// ============================================================================
kernel void kernel_name(
device const float* input [[buffer(0)]],
device float* output [[buffer(1)]],
constant uint& size [[buffer(2)]],
uint tid [[thread_position_in_grid]]
) {
// ... implementation
}Naming: snake_case for all identifiers (functions, variables, parameters)
Inline Functions:
inline float bf16_to_f32(uint16_t bf16) {
return as_type<float>(uint(bf16) << 16);
}Style: PEP 8 compliant, snake_case for functions/variables
Docstrings: Triple quotes, one-line for simple functions
def parse_layers(spec):
"""Parse layer specification like '0-4' or '0,5,10'."""
if spec is None or spec == 'all':
return list(range(60))
# ...Imports: Standard library first, then third-party
import argparse
import json
import os
import sys
import pandas as pd
import matplotlib.pyplot as pltThe launcher (flashchat) and any helper script that needs to create a venv or run pip install must source lib/python.sh and call flashchat_python_bin. Never call python3 directly in shell. Reasons:
- macOS 26 ships
/usr/bin/python3 = 3.9from CommandLineTools. Thatpython3shadows any Homebrew Python the user has installed (/opt/homebrew/bin/python3is later on PATH for shells that have runbrew shellenv, but a fresh terminal or a non-interactive script can still hit CLT first). python3 -m venvinherits the parent interpreter. A venv is a thin wrapper around an existing Python binary — there is no "use a different Python" knob. If youpython3 -m venvagainst CLT 3.9, your venv is 3.9, full stop.modelmgr/uses PEP 604 union syntax unconditionally (str | None,dict | None, ~50+ sites acrossartifacts.py,manifest.py,migrate.py,configfile.py,status.py, etc.). 3.9 cannot parse these at import time, so the firstmodelmgrimport raisesTypeError: unsupported operand type(s) for |: 'type' and 'NoneType'and the user gets a "Setup incomplete" screen for what looks like a setup bug.
flashchat_python_bin resolves in this priority:
$FLASHCHAT_PYTHON— explicit override (path is verified executable and meets the version floor)./opt/homebrew/bin/python3— Apple Silicon Homebrew. Probed, must report ≥ 3.10./usr/local/bin/python3— Intel Homebrew legacy prefix. Probed, must report ≥ 3.10.python3from PATH — last resort, must report ≥ 3.10.
If no candidate satisfies the floor, the helper exits non-zero with a clear actionable message ("brew install python" or "set FLASHCHAT_PYTHON"). It never modifies the user's PATH, default python3, or any system-level state.
Minimum supported Python: 3.10. PEP 604 is the floor. Do not introduce match statements or PEP 695 type aliases without also bumping FLASHCHAT_MIN_PYTHON_MINOR in lib/python.sh.
Stale-venv handling. ensure_venv in the launcher checks whether the existing metal_infer/.venv/bin/python meets the floor; if it doesn't (e.g., the venv was created with CLT 3.9 from a previous flashchat run), the venv is recreated silently using the resolver's chosen interpreter. A single Recreating project venv (existing interpreter below Python 3.10). line is logged. This is intentional: regenerating metal_infer/.venv/ is regenerating derived state, not user data — no consent screen needed.
Testing the resolver in isolation:
# Default (PATH has whatever it has)
source lib/python.sh && flashchat_python_bin
# Simulate a CLT-only PATH — should still find Homebrew Python
PATH="/usr/bin:/bin:/usr/sbin:/sbin" bash -c 'source lib/python.sh && flashchat_python_bin'
# Force-fail: pin to CLT 3.9, expect non-zero exit and an actionable error
FLASHCHAT_PYTHON=/usr/bin/python3 bash -c 'source lib/python.sh && flashchat_python_bin'- CMD1: Attention projections + delta-net (GPU)
- CPU: Flush results, softmax + topK routing
- SSD: Parallel pread K=4 experts (2.4ms)
- CMD2: o_proj + norm + routing (GPU)
- CMD3: Expert forward + combine + norm (GPU, deferred)
- MLX affine 4-bit, group_size=64
- Weights: uint32 (8 x 4-bit values packed)
- Per-group scale and bias in bfloat16
- Dequant:
value = uint4_val * scale + bias
- Non-expert weights: mmap'd (5.5GB, read-only)
- Expert weights: Streamed from SSD on demand
- Metal scratch: ~200MB
- Trust OS page cache for expert data (no custom LRU)
- Deferred GPU command submission for pipeline overlap
- FMA-optimized dequant kernels
- BLAS (Accelerate) for linear attention
- Parallel
pread()for expert I/O