An exhaustive, numbered list of what runs on a gaming PC, on an AI-creator PC, and with the NYMPH AX1 Premium card. Every figure carries its provenance tag — where measurement ends, marketing begins, and that line is marked here.
REAL = compiles / patent filed. MED = measured, aggregated from the community (llama.cpp, Unsloth, public benchmarks). EST = roofline calibrated against measurements. PROJ = NYMPH projection — zero tok/s measured on Axera yet; ceiling/simulation until the card ships ($199). All figures are ±30% bands (they vary with context, quantization and cooling).
A PC already has five compute engines that were never introduced to each other. The NYMPH firmware routes every operation — language, vision, voice, retrieval, draft — to the optimal engine, and frees or recruits your GPU as needed.
Commercial rule: parallelism is guaranteed (each chip runs its own thing). Splitting one model across 2 chips (sharding) is roadmap — Pulsar2 allows it, pending on-silicon validation. Never promise "one 48-TOPS brain" if the SDK delivers "two 24-TOPS brains."
The honest floor: this is community-measured, free and local, without buying anything. Two machine profiles.
| Model | Type | Gamer 8 GB | Creator 16 GB | Source |
|---|---|---|---|---|
| Qwen3 8B | dense | 35–45 | 70–90 | MED |
| Gemma 4 12B | dense · multimodal | 20–30 | 45–60 | MED |
| Qwen3 14B | dense | 12–18 | 40–55 | MED |
| Mistral / Magistral 24B | dense | 5–8 | 28–40 | MED |
| Qwen3 32B | dense | 3–5 | 10–15 | MED |
| Llama 3.3 70B | dense | — doesn't fit | 3–5 | MED |
| Qwen3.6 35B-A3B | MoE (3B act.) | 20–25 | 70–90 | MED |
| Gemma 4 26B-A4B | MoE (4B act.) | 18–22 | 55–75 | EST |
| gpt-oss-20B | MoE | 25–35 | 60–80 | EST |
| Qwen3-Next 80B-A3B | MoE (3B act.) | 5–8 | 12–18 | EST |
| gpt-oss-120B | MoE (~5B act.) | — no | 10–15 | EST |
The golden rule: in dense, the usable ceiling is what fits in VRAM — ~14B on the Gamer, ~24B on the Creator; crossing it drops you to 1–5 tok/s. In MoE the active count rules, not the total: a 35B-A3B runs faster than a 14B dense. "80B at 30 tok/s on consumer" does not exist — a usable 80B gives ~15.
| Model | Params | Gamer 8 GB | Creator 16 GB | Best at |
|---|---|---|---|---|
| SD 1.5 | 0.86B | ~2–4 s | <2 s | huge LoRA library |
| SDXL 1.0 (+Pony/Illustrious) | 2.6B | ~15–25 s | ~6–10 s | deepest LoRA ecosystem |
| SD 3.5 Large | 8B | tight/no | ~12–18 s | photorealism |
| FLUX.1 dev | 12B DiT | ~40–90 s | ~12–20 s | best prompt adherence |
| FLUX.1 schnell (4 steps) | 12B | ~10–20 s | ~2–4 s | Apache 2.0, commercial |
| Z-Image Turbo (8 steps) | 6B | ~10–15 s | ~3–5 s | commercial, bilingual, fast |
| Qwen-Image | 20B MMDiT | slow | ~30–60 s | best legible in-image text |
| FLUX.2 dev | 32B | — no | ~25–40 s | top quality, heavy |
Image is compute-bound: the Gamer runs almost the same models, just slower; VRAM only blocks the biggest ones. Standard tool: ComfyUI.
| Model | Params | Gamer 8 GB | Creator 16 GB | Note |
|---|---|---|---|---|
| LTX-Video / LTX-2 | 13–19B | ~7–9 min | ~30–90 s | speed champion; LTX-2 = synced audio+video |
| Wan 2.2 TI2V-5B | 5B | Wan 1.3B only | ~2–4 min | quality leader, Apache 2.0 |
| Wan 2.2 A14B | 27B / 14B act. | no | ~4–9 min | higher quality, heavier |
| HunyuanVideo 1.5 | 8.3B | no | ~75 s–6 min | best faces / human motion |
| CogVideoX-5B | 5B | tight | ~3–5 min | 480p, image-to-video |
Local video is minutes-per-clip except LTX — for iteration and B-roll, not clip-after-clip like the cloud. The 8 GB Gamer is essentially LTX-only.
| Model | Task | Gamer 8 GB | Creator 16 GB | Note |
|---|---|---|---|---|
| Whisper large-v3 | speech→text | 8–12× | 15–20× | 99 languages, top accuracy |
| large-v3-turbo | speech→text | ~15× | ~30× | ideal for live voice |
| Kokoro-82M | text→speech | 15–210× | 15–210× | #1 TTS Arena, Apache 2.0, runs on CPU |
| XTTS v2 / Chatterbox | text→speech (clone) | ~3× | ~3× | voice cloning from 6 s |
2× Axera AX8850 · 48 TOPS NPU (+6 orchestration) · 16 GB of AI memory · 64 GB persistent · ~16 W. The firmware already compiles and passes smoke 9/9 in simulation. Zero tok/s measured on Axera — every speed here is ceiling/projection until physical silicon.
Close, reboot, return — it greets you by name. State Capsules (USPTO 63/901,576). REAL (firmware in sim).
The agent + RAG + senses run on the card at ~16 W while you game or render with the GPU 100% free. REAL (architecture).
A 26B no 8–16 GB gaming GPU can load — it fits on the card. REAL (measured 0.54 byte/param law).
Isolated state, every inference audited <2 ms, zero bytes to the cloud. REAL (Governance in firmware).
| Model | Type | tok/s (real.) | Fits? | Tag |
|---|---|---|---|---|
| Qwen3-0.6B | dense | ~20–40 | 1 module | ROOFLINE |
| ~4B dense | dense | ~9 | 1 module | ROOFLINE |
| 7B distill | dense | ~5–9 | 1 module | ROOFLINE / 3P (~13?) |
| Qwen3-8B | dense | ~4–8 | 1 module (compiles REAL, 4.43 GB) | ROOFLINE |
| Qwen3-14B | dense | ~2.5–4 | tight | ROOFLINE |
| Gemma / Qwen 26B | dense | ~1.5–2.5 | 2 modules (16 GB) | capacity, not speed |
| Qwen3-32B | dense | ~1.2–2 | 2 modules, tight | background only |
| Qwen3-30B-A3B | MoE (3B act.) | ~9–11 | 2 mod · or stream in 8 GB (7.5–9.3) | ROOFLINE / PROJ |
| Gemma 26B-A4B (Trinity's brain) | MoE (4B act.) | ~7–9 | 2 modules (14 GB) | ROOFLINE |
| Mamba ~2B (native SSM) | SSM · no KV | ~15–18 | compiles REAL · speed ROOFLINE | |
| Mamba ~4B | SSM · no KV | ~9 | 1 module | ROOFLINE |
| Draft Mamba 300–500M | spec-decode draft | tens | 1 module | ROOFLINE |
| Spec-decode (draft + 26B) | draft+verify | ~3 | 2 modules | PROJECTED (~2× verifier) |
The hard read: large dense = capacity, not speed (26B ~1.5–2.5 tok/s, for background agents). The truly fast = MoE + Mamba + spec-decode — that's where the projected double digits are. The defensible headline: "runs 26–30B MoE no gaming GPU can load, at ~16 W, always-on; targets ~9–11 tok/s — pending on-silicon measurement."
| Workload | Status on Axera | Tag |
|---|---|---|
| Temporal video (Conv3D) | unlocked — beyond typical edge NPUs | compiles REAL |
| Image (SDXL/FLUX VAE + UNet/DiT) | operators map, viable with 16 GB | compiles REAL |
| Audio (Whisper / CosyVoice2 / ASR) | out of the box, ~3 W voice assistant | compiles REAL |
| Multimodal / VLM (Qwen-VL / OCR) | out of the box — VQA, docs, screen-reading | compiles REAL |
| RAG / embeddings | whole pipeline on-card, 100% NPU | compiles REAL |
NYMPH doesn't change what your GPU already does — it keeps it and adds what the GPU can't. That's why "with NYMPH" retains everything from "without NYMPH" and frees the GPU on top.
The sentence that sums up both: the PC alone is already fast — NYMPH doesn't compete there. It adds what no GPU gives itself: persistence, always-on with a free GPU, an extra brain above your VRAM ceiling, and auditable isolation. Like in the 90s your PC had no sound → Sound Blaster; in 2026 it can't run AI without sacrificing everything else → NYMPH.
Decode is memory-bound: speed is set by the bandwidth of the tier where the active weight lives. The pooled hybrid adds tiers up to ~50–100 GB (roadmap).
Host↔card bottleneck: PCIe Gen2 (~0.9 GB/s effective). That's why expert-streaming and the pooled hybrid live or die on cache hit-rate, not raw bandwidth.
| Capability | PC alone (8–16 GB) | PC + AX1 Premium |
|---|---|---|
| Fast local chat (35B-A3B MoE) | ✅ 20–90 tok/s MED | ✅ same (the GPU does it) |
| Memory surviving reboot | ❌ | ✅ REAL (sim) |
| 24/7 AI with a free GPU | ❌ GPU hijacked | ✅ ~16 W |
| Extra resident 26–30B model | ❌ doesn't fit | ✅ fits REAL · ~9–11 PROJ |
| Permanent vision + voice without GPU | ❌ | ✅ compiles 100% NPU REAL |
| Deterministic per-inference audit | ❌ | ✅ |
| Structural privacy (isolated domain) | ❌ | ✅ |
| Cost | $0 | $1,190 |
The PC alone is already good at speed; NYMPH doesn't compete there. It adds the four things the PC can't give itself: persistence, always-on, extra model capacity, and auditable isolation. vs buying a $4,000 AI machine (DGX Spark): that wins raw tok/s and model ceiling, but it's another machine; AX1 empowers the one you already own and keeps your GPU free.
Method: "without NYMPH" figures = community measurements/aggregates (llama.cpp, Unsloth, public benchmarks) or roofline calibrated against them. "With NYMPH" figures = bandwidth ceiling (roofline over spec) or own simulation projection — no tok/s is measured on a physical AX-M1; what compiles and how much it occupies IS measured in the Pulsar2 compiler. All are ±30% bands and vary with host, context, quantization and cooling. The nymph-firmware compiles and passes smoke 9/9 in simulation; silicon backends are 1:1 drop-ins when the board ships. Cited patents = USPTO provisionals (require conversion to non-provisional/PCT). NYMPH is a trademark of Punky Tiger Labs, Inc. AX8850 and AX-M1 are products of Axera Semiconductor; AICore AX-M1 is by Radxa. © 2026 Punky Tiger Labs, Inc.