Compatibility · spec sheet

Every model you can run.

An exhaustive, numbered list of what runs on a gaming PC, on an AI-creator PC, and with the NYMPH AX1 Premium card. Every figure carries its provenance tag — where measurement ends, marketing begins, and that line is marked here.

REAL MED · community EST · roofline PROJ · unmeasured

REAL = compiles / patent filed.   MED = measured, aggregated from the community (llama.cpp, Unsloth, public benchmarks).   EST = roofline calibrated against measurements.   PROJ = NYMPH projection — zero tok/s measured on Axera yet; ceiling/simulation until the card ships ($199). All figures are ±30% bands (they vary with context, quantization and cooling).

The principle

Every job, to its own engine.

A PC already has five compute engines that were never introduced to each other. The NYMPH firmware routes every operation — language, vision, voice, retrieval, draft — to the optimal engine, and frees or recruits your GPU as needed.

NYMPH ROUTER (NUAIP)
routes by workload type · load-aware · with fallback
Host GPU
large LLM, image, video
8–16 GB · 350–900 GB/s
2× NPU AX8850
resident LLM, vision, audio, embeddings
48 TOPS · 16 GB · ~16 W
RK3588 (orchestrator)
router, safety-gate, draft ≤2B
6 TOPS · Linux on-card
Host CPU + RAM
autoregressive loop, KV, cold experts
32–64 GB · 45–90 GB/s
NVMe + NAND
persistent cognitive memory
64 GB state · survives reboot

Commercial rule: parallelism is guaranteed (each chip runs its own thing). Splitting one model across 2 chips (sharding) is roadmap — Pulsar2 allows it, pending on-silicon validation. Never promise "one 48-TOPS brain" if the SDK delivers "two 24-TOPS brains."

Without NYMPH · what already runs today

Your PC, as is, running local AI.

The honest floor: this is community-measured, free and local, without buying anything. Two machine profiles.

🎮 GAMING PC · 8 GB
  • GPU 8 GB VRAM · ~350 GB/s (RTX 4060 / 3060 Ti)
  • RAM 32 GB DDR4-3200 · ~45 GB/s
  • Link PCIe 3.0/4.0
🎨 AI-CREATOR PC · 16 GB
  • GPU 16 GB VRAM · ~900 GB/s (RTX 4080 / 5070 Ti)
  • RAM 64 GB DDR5-6000 · ~90 GB/s
  • Link PCIe 4.0 x16 · ~31 GB/s
LANGUAGE (LLM) — generation tok/s, Q4, batch 1 · higher = better
ModelTypeGamer 8 GBCreator 16 GBSource
Qwen3 8Bdense35–4570–90MED
Gemma 4 12Bdense · multimodal20–3045–60MED
Qwen3 14Bdense12–1840–55MED
Mistral / Magistral 24Bdense5–828–40MED
Qwen3 32Bdense3–510–15MED
Llama 3.3 70Bdense— doesn't fit3–5MED
Qwen3.6 35B-A3BMoE (3B act.)20–2570–90MED
Gemma 4 26B-A4BMoE (4B act.)18–2255–75EST
gpt-oss-20BMoE25–3560–80EST
Qwen3-Next 80B-A3BMoE (3B act.)5–812–18EST
gpt-oss-120BMoE (~5B act.)— no10–15EST

The golden rule: in dense, the usable ceiling is what fits in VRAM — ~14B on the Gamer, ~24B on the Creator; crossing it drops you to 1–5 tok/s. In MoE the active count rules, not the total: a 35B-A3B runs faster than a 14B dense. "80B at 30 tok/s on consumer" does not exist — a usable 80B gives ~15.

IMAGE — seconds per 1024×1024 image · lower = better
ModelParamsGamer 8 GBCreator 16 GBBest at
SD 1.50.86B~2–4 s<2 shuge LoRA library
SDXL 1.0 (+Pony/Illustrious)2.6B~15–25 s~6–10 sdeepest LoRA ecosystem
SD 3.5 Large8Btight/no~12–18 sphotorealism
FLUX.1 dev12B DiT~40–90 s~12–20 sbest prompt adherence
FLUX.1 schnell (4 steps)12B~10–20 s~2–4 sApache 2.0, commercial
Z-Image Turbo (8 steps)6B~10–15 s~3–5 scommercial, bilingual, fast
Qwen-Image20B MMDiTslow~30–60 sbest legible in-image text
FLUX.2 dev32B— no~25–40 stop quality, heavy

Image is compute-bound: the Gamer runs almost the same models, just slower; VRAM only blocks the biggest ones. Standard tool: ComfyUI.

VIDEO — minutes per ~5 s clip (720p) · lower = better
ModelParamsGamer 8 GBCreator 16 GBNote
LTX-Video / LTX-213–19B~7–9 min~30–90 sspeed champion; LTX-2 = synced audio+video
Wan 2.2 TI2V-5B5BWan 1.3B only~2–4 minquality leader, Apache 2.0
Wan 2.2 A14B27B / 14B act.no~4–9 minhigher quality, heavier
HunyuanVideo 1.58.3Bno~75 s–6 minbest faces / human motion
CogVideoX-5B5Btight~3–5 min480p, image-to-video

Local video is minutes-per-clip except LTX — for iteration and B-roll, not clip-after-clip like the cloud. The 8 GB Gamer is essentially LTX-only.

AUDIO — × real-time · higher = better · never the bottleneck, runs in parallel with the LLM in ~2 GB
ModelTaskGamer 8 GBCreator 16 GBNote
Whisper large-v3speech→text8–12×15–20×99 languages, top accuracy
large-v3-turbospeech→text~15×~30×ideal for live voice
Kokoro-82Mtext→speech15–210×15–210×#1 TTS Arena, Apache 2.0, runs on CPU
XTTS v2 / Chatterboxtext→speech (clone)~3×~3×voice cloning from 6 s
With NYMPH AX1 Premium · what your GPU can't do

The card doesn't compete on speed. It adds what's missing.

2× Axera AX8850 · 48 TOPS NPU (+6 orchestration) · 16 GB of AI memory · 64 GB persistent · ~16 W. The firmware already compiles and passes smoke 9/9 in simulation. Zero tok/s measured on Axera — every speed here is ceiling/projection until physical silicon.

Remembers after reboot

Close, reboot, return — it greets you by name. State Capsules (USPTO 63/901,576). REAL (firmware in sim).

24/7 without touching your GPU

The agent + RAG + senses run on the card at ~16 W while you game or render with the GPU 100% free. REAL (architecture).

Models above your VRAM ceiling

A 26B no 8–16 GB gaming GPU can load — it fits on the card. REAL (measured 0.54 byte/param law).

Structural privacy

Isolated state, every inference audited <2 ms, zero bytes to the cloud. REAL (Governance in firmware).

LANGUAGE ON THE CARD — realistic tok/s (band @~20 GB/s effective) · INT4 · ALL PROJECTED
ModelTypetok/s (real.)Fits?Tag
Qwen3-0.6Bdense~20–401 moduleROOFLINE
~4B densedense~91 moduleROOFLINE
7B distilldense~5–91 moduleROOFLINE / 3P (~13?)
Qwen3-8Bdense~4–81 module (compiles REAL, 4.43 GB)ROOFLINE
Qwen3-14Bdense~2.5–4tightROOFLINE
Gemma / Qwen 26Bdense~1.5–2.52 modules (16 GB)capacity, not speed
Qwen3-32Bdense~1.2–22 modules, tightbackground only
Qwen3-30B-A3BMoE (3B act.)~9–112 mod · or stream in 8 GB (7.5–9.3)ROOFLINE / PROJ
Gemma 26B-A4B (Trinity's brain)MoE (4B act.)~7–92 modules (14 GB)ROOFLINE
Mamba ~2B (native SSM)SSM · no KV~15–18compiles REAL · speed ROOFLINE
Mamba ~4BSSM · no KV~91 moduleROOFLINE
Draft Mamba 300–500Mspec-decode drafttens1 moduleROOFLINE
Spec-decode (draft + 26B)draft+verify~32 modulesPROJECTED (~2× verifier)

The hard read: large dense = capacity, not speed (26B ~1.5–2.5 tok/s, for background agents). The truly fast = MoE + Mamba + spec-decode — that's where the projected double digits are. The defensible headline: "runs 26–30B MoE no gaming GPU can load, at ~16 W, always-on; targets ~9–11 tok/s — pending on-silicon measurement."

OTHER MODALITIES ON THE CARD — compile 100% NPU (25/26 operators) · speed = board
WorkloadStatus on AxeraTag
Temporal video (Conv3D)unlocked — beyond typical edge NPUscompiles REAL
Image (SDXL/FLUX VAE + UNet/DiT)operators map, viable with 16 GBcompiles REAL
Audio (Whisper / CosyVoice2 / ASR)out of the box, ~3 W voice assistantcompiles REAL
Multimodal / VLM (Qwen-VL / OCR)out of the box — VQA, docs, screen-readingcompiles REAL
RAG / embeddingswhole pipeline on-card, 100% NPUcompiles REAL
Side by side · your machine, without and with the card

The same PC, before and after NYMPH.

NYMPH doesn't change what your GPU already does — it keeps it and adds what the GPU can't. That's why "with NYMPH" retains everything from "without NYMPH" and frees the GPU on top.

🎮 GAMING PC · 8 GB VRAM · 32 GB RAM
Without NYMPH
  • LLM: 35B-A3B MoE at 20–25 tok/s MED — but while it runs, it hijacks your GPU
  • Image: SDXL ~15–25 s · FLUX schnell ~10–20 s
  • Video: LTX only, ~7–9 min/clip
  • Voice: Whisper turbo ~15× real time
  • Memory: ✕ starts from zero every session
  • 24/7: ✕ can't game while the AI uses the GPU
  • Model > 14B dense: ✕ 1–5 tok/s, unusable
With NYMPH AX1 Premium
  • Everything on the left STILL WORKS — the GPU does it — and now the GPU is 100% free to game
  • + Resident 26–30B MoE brain on the card at ~9–11 tok/sPROJ — that your GPU can't load
  • + Memory that survives reboot REAL sim
  • + 24/7 agent at ~16 W while you game
  • + Permanent vision and voice on the NPU, not touching the GPU
  • + Local private code / RAG — never leaves the machine
  • + Every inference audited <2 ms
🎨 AI-CREATOR PC · 16 GB VRAM · 64 GB RAM
Without NYMPH
  • LLM: 35B-A3B MoE at 60–90 tok/s MED, or an 80B-A3B at ~15 — but they take the whole VRAM
  • Image: FLUX.1 dev ~12–20 s · SDXL ~6–10 s
  • Video: LTX ~30–90 s · Wan 2.2 ~2–4 min
  • Voice: Whisper large-v3 ~15–20× real time
  • Memory: ✕ no persistence across reboots
  • 24/7 with a free GPU: ✕ the LLM or the render, not both
  • Model > 24B dense: ✕ drops to slow offload
With NYMPH AX1 Premium
  • Everything on the left STILL WORKS, and now the GPU is 100% free for render/heavy video
  • + 2nd resident 26–30B brain in parallel — doesn't share your VRAM PROJ ~9–11
  • + Speculative decoding: draft on the NPU + verify on your GPU → compute that adds upPROJ
  • + Two pipelines at once: permanent vision on one module, LLM+embeddings on the other
  • + Full Trinity / PPAi on-card: identity, episodic/emotional memory, daemons REAL sim
  • + Persistent, auditable AI node — 26B quality at zero marginal cost per token

The sentence that sums up both: the PC alone is already fast — NYMPH doesn't compete there. It adds what no GPU gives itself: persistence, always-on with a free GPU, an extra brain above your VRAM ceiling, and auditable isolation. Like in the 90s your PC had no sound → Sound Blaster; in 2026 it can't run AI without sacrificing everything else → NYMPH.

Memory hierarchy

Where each weight lives.

Decode is memory-bound: speed is set by the bandwidth of the tier where the active weight lives. The pooled hybrid adds tiers up to ~50–100 GB (roadmap).

GPU VRAM
8–16 GB · 350–900 GB/s — the fastest
AX card RAM
16 GB · ~34 GB/s (LPDDR4x) — the resident brain
NVMe / NAND
64 GB · persistent state, survives reboot
Host RAM
32–64 GB · 45–90 GB/s — cold experts, KV-pin
Host disk
TB · full model for expert-streaming

Host↔card bottleneck: PCIe Gen2 (~0.9 GB/s effective). That's why expert-streaming and the pooled hybrid live or die on cache hit-rate, not raw bandwidth.

Summary

PC alone vs PC + NYMPH AX1 Premium.

CapabilityPC alone (8–16 GB)PC + AX1 Premium
Fast local chat (35B-A3B MoE)✅ 20–90 tok/s MED✅ same (the GPU does it)
Memory surviving rebootREAL (sim)
24/7 AI with a free GPU❌ GPU hijacked✅ ~16 W
Extra resident 26–30B model❌ doesn't fit✅ fits REAL · ~9–11 PROJ
Permanent vision + voice without GPUcompiles 100% NPU REAL
Deterministic per-inference audit
Structural privacy (isolated domain)
Cost$0$1,190

The PC alone is already good at speed; NYMPH doesn't compete there. It adds the four things the PC can't give itself: persistence, always-on, extra model capacity, and auditable isolation. vs buying a $4,000 AI machine (DGX Spark): that wins raw tok/s and model ceiling, but it's another machine; AX1 empowers the one you already own and keeps your GPU free.

Spec sheet · the year's laws

What governs every number.

1Decode = memory-bound. tok/s ≈ effective bandwidth ÷ active bytes per token. Any optimization not attacking data movement is secondary.
2INT4 law ≈ 0.54 byte/param (INT8 ≈ 1.10), measured in Pulsar2. With it: a 26B ≈ 14 GB → fits in 16 GB. (Small models rise to ~0.72 from embed/vocab.)
3MoE multiplies speed ~5–7× vs same-size dense — reads only the active experts (3–4B/token). Small hardware's sweet spot is high sparsity (A3B/A4B).
4Large dense = capacity, not speed (26B ≈ 1.5–2.5 tok/s). Position as a background agent, never live chat.
5Bandwidth is bottleneck #1, not TOPS. The GPU in LLM inference runs at 15–40% — it's waiting for data, not computing.
6Measure before publishing. No tok/s as fact without on-silicon measurement. The single external datapoint (~13 tok/s on 7B) doesn't reconcile with bandwidth physics (34 GB/s) — reason #1 to measure.

Method: "without NYMPH" figures = community measurements/aggregates (llama.cpp, Unsloth, public benchmarks) or roofline calibrated against them. "With NYMPH" figures = bandwidth ceiling (roofline over spec) or own simulation projection — no tok/s is measured on a physical AX-M1; what compiles and how much it occupies IS measured in the Pulsar2 compiler. All are ±30% bands and vary with host, context, quantization and cooling. The nymph-firmware compiles and passes smoke 9/9 in simulation; silicon backends are 1:1 drop-ins when the board ships. Cited patents = USPTO provisionals (require conversion to non-provisional/PCT). NYMPH is a trademark of Punky Tiger Labs, Inc. AX8850 and AX-M1 are products of Axera Semiconductor; AICore AX-M1 is by Radxa. © 2026 Punky Tiger Labs, Inc.