Punky Tiger Labs · SUPERION AX2
A card that keeps whole models loaded in your PC — so agents can work all night, remember what they did, and never send a byte to anyone else's cloud. You buy it once.
One-time payment · No subscription · No per-token bill
The problem
Chat apps are great at answering. They aren't built to hold a role, work overnight, or remember you tomorrow. Cloud agents send everything you touch to somebody else's machine and bill you by the token. And the industry's only answer is to sell you another graphics card.
SUPERION AX2 says the opposite. Your processor works. Your memory works. What was missing is a place where the model can live — memory that belongs to the AI and never gives it back.
Two modes, one switch
Factory mode
The card, your graphics card, your processor and your system memory all compute at once — so a model far larger than your GPU could hold on its own answers faster than you can read.
Resident mode
The card alone. Your GPU and processor stay free for your own work while models stay loaded, agents keep their memory, and voice, vision and search run in the background.
Proof, not a promise
Trinity runs her own radio station — picking the records, reading the news, thinking out loud between songs, live and unscripted, around the clock. She is the same kind of persistent mind SUPERION AX2 is built to keep resident on your desk. Press play and listen to one working right now.
Trinity is our own persistent AI, running on our hardware. Nothing you hear is pre-recorded.
Compare
We'll be straight with you: if you want the single most capable model on earth this afternoon, that model is in a data centre and it isn't ours. What a data centre cannot give you is a mind that stays, on hardware you own, with no meter running.
| Cloud API | Agent box | SUPERION AX2 | |
|---|---|---|---|
| Where the model runs | Someone else's cloud | Someone else's cloud | Your machine |
| Where your data goes | Out | Out | Nowhere |
| Works with no internet | No | No | Yes |
| Cost per token | Metered, forever | Metered, forever | None, ever |
| Rate limits | Yes | Yes, upstream | None |
| Agents keep memory after shutdown | No | Depends on the vendor | Yes |
| Frontier-scale models | Yes — best in class | Yes, via the cloud | No — up to what the card holds |
| What you pay | Every month | Hardware + usage | Once |
An "agent box" is a small always-on computer that runs agent software but still calls a cloud model to think. It solves where the agent runs. SUPERION AX2 solves where the model lives.
Agents
Three models stay resident at once — one per engine. Agents hold their own context, sleep when idle, and wake with everything they knew still intact. Close the lid, come back tomorrow, and the one that was watching your inbox is still the one that was watching your inbox.
It speaks the API your tools already speak, and ships an MCP server — so the assistant you use today can reach it without you rewriting anything.

Pre-order
Identical engines, identical memory, identical software rights. Storage does not change how fast it thinks — only how much it can keep.
Full technical report, model catalogue and measurement protocol available on request.
FAQ
An add-in card for a desktop PC. It carries two AI accelerators and an orchestrator with their own 22.7 GB of memory, plus a position for an NVMe drive. Models load into that memory and stay there — so the AI is a resident of your machine rather than a request to somebody's server.
No, and it isn't meant to. Your GPU keeps doing what it's good at. AX2 adds memory and compute that belong to the AI, and can either work together with your GPU on one big model, or work entirely on its own so your GPU stays free for your games and your work.
That's the point of resident mode. The card carries the agents by itself; your processor and graphics card are untouched.
Yes. The models are on the card and the drive. Nothing about answering you requires a network.
No. AX2 speaks a compatible API and ships an MCP server, so the assistant you already use can reach it and hand it the work that should stay home. Frontier models in the cloud are still better at the hardest single questions. AX2 is for everything that should be permanent, private, or always running.
Three models stay resident at once, one per engine, with many agents sharing them and sleeping to NVMe when idle. The exact live count depends on the models you choose and how hard you push them — the technical report gives the measured numbers model by model.
Not through us. There is no account, no telemetry and no cloud call in the path between you and a resident model.
A desktop PC with a free PCIe x16 mechanical slot, two slots of physical clearance, and enough power headroom for a 50 W card. Windows and Linux.
Because we build them in batches. The pre-order tells us how many to build and holds your unit in the first one. The deposit is refundable until we charge the balance.
Buy the machine once. Keep everything it learns.