16 GB of dedicated AI memory runs a 26B-class brain no gaming GPU can even load — text, image, video, audio and vision, native on silicon. Private, always on.
Every modality, on one private card. Nothing sent to the cloud.

16 GB of AI memory fits a professional model your GPU can't load — private, on your machine.
Text, image, video, audio and vision — all native on the silicon, out of the box.
Audio, memory, retrieval and perception move to dedicated silicon. Your GPU stays 100% yours.
64 GB persistent state. Your AI remembers — it never starts from zero.
Production-grade local AI, including NYMPH AI — plus an open toolchain to run and supercharge any open-source model. Bigger transformers run on your GPU under orchestration; the rest runs on the card.
A patent-pending orchestration layer coordinates the card and your host —CPU, RAM, GPU— as one system. NYMPH carries the resident brain, memory, retrieval and perception; your GPU keeps 100% of its power for what it was bought for.
One OpenAI-compatible local endpoint and a native MCP server. Point your tools at NYMPH and they just work — now private, persistent and offline.
What NYMPH adds to any of them: persistent memory · private local RAG · offline fallback · a freed GPU. It also makes any open-source AI run better — local, private and always on.
Your whole repo lives on the card: persistent project memory, private local RAG, decisions remembered, commands policy-gated. It stops re-reading your codebase every session — and your code never leaves the machine.
On-card skill routing, persistent agent memory and a safety gate. Your agent remembers across runs, validates its own actions, and keeps working offline — at near-zero cost.
Point Codex at a local OpenAI-compatible endpoint: private code and context, no per-token bill, and an offline fallback when you want it. Same workflow, on your hardware.
On-device AI for the browser and apps: local transcription, summarization, search and memory — fast, private, nothing sent out.
Give ChatGPT, Claude or any model a private local memory, an offline fallback, and lower cost. Your assistant stops forgetting.
A resident coding engineer via MCP: your repo indexed on-card, private RAG, every command policy-gated in under 2 ms. Code never leaves your machine.
Real NPCs that come alive — characters that see, hear, remember you and react in real time, on dedicated silicon so your GPU keeps 100% for the game. Build AI-native games, live voice and overlays — all local.

Pre-orders open August 5. Register your interest — reserve your place with a fully refundable deposit when reservations open.
Pre-orders open August 5Performance and capacity figures are based on component-level benchmarks, the Axera Pulsar2 toolchain and pre-production engineering data; system performance varies with host and workload. Model-fit figures (INT4) are compile-validated; on-device throughput is pre-production and not yet field-measured. Target MSRP; final pricing set at general availability. Specifications subject to change. NYMPH is a trademark of Punky Tiger Labs, Inc. AX8850 and AX-M1 are products of Axera Semiconductor; AICore AX-M1 is a product of Radxa. RK3588 is a product of Rockchip. © 2026 Punky Tiger Labs, Inc.