◆ edge AI terminal · desktop

Run models on your desk, not someone's cloud.

gpuvera is a palm-sized inference terminal. Plug it in, load a model, and prompt it over your own network — no accounts, no data leaving the room, no monthly fee to keep it working.

Private by default · Works fully offline · Yours to own — no subscription

gpuvera desktop terminal on a desk, three-quarter view showing USB-C, USB-A, Ethernet and power ports with a machined cooling grille — concept render, pre-production.

Concept render — pre-production. Not a photo of a shipping unit.

Why local

Your model, your data, your electricity.

Cloud inference means your prompts, your documents, and your usage all live on someone else's servers. gpuvera keeps the whole loop on your desk.

🔒

Private by default

Prompts and files never leave the device unless you send them. Nothing is logged to a vendor account, because there is no vendor account.

Works offline

Once a model is loaded, gpuvera runs with the network unplugged. Planes, labs, air-gapped rooms, spotty Wi-Fi — it doesn't care.

You own the compute

Buy it once. No per-token billing, no "pro" tier, no feature that stops working when a subscription lapses. The hardware is yours.

How it works

Three steps, then it's just there.

1 · Plug it in

USB-C power in, Ethernet or Wi-Fi for setup. It sits next to your monitor, not in a server rack — palm-sized, passively-then-actively cooled, quiet on a desk.

Footprint: fits in one hand · desk-quiet cooling*

Closeup of gpuvera rear I/O: two USB-C ports, Ethernet, DC power, hex screws and a machined cooling grille — concept render, pre-production.

2 · Load a model

Open the local console in your browser and pull an open-weight model — chat, code, or vision. gpuvera stores it on-device so it's ready next time without another download.

Runs common open-weight models in the [ __ ]B class*

gpuvera terminal beside a laptop running a local model console and dashboard on a desk — concept render, pre-production.

3 · Prompt it — from anywhere on your LAN

Your laptop, phone, or another workstation talks to gpuvera over your own network. One device serves the whole room, and the traffic never touches the public internet.

Local HTTP API · OpenAI-compatible endpoint*

gpuvera terminal next to an open human hand for scale, roughly palm-sized — concept render, pre-production.
Straight talk

What we won't pretend.

A desk terminal is not a datacenter. We'd rather you know the trade before you buy.

It runs small-to-mid models

Frontier 400B+ models want a cluster. gpuvera targets the open-weight models people actually run locally — sized to fit its memory, stated plainly on the specs page.

Tokens/sec, measured honestly

Every throughput number will ship with the model, quantization, and prompt length it was measured at. No cherry-picked hero figure.

It's early

These are pre-production units and the specs below are targets, marked as such. When numbers are final and verified, we'll say so — and note anything that changed.

Reserve a unit

Want one on your desk?

Join the waitlist for pricing, ship window, and refund terms as they firm up. No charge to reserve, no spam.

Join the waitlist →