tokenquestBY KUBESIMPLIFY
YOUR ADVENTURE STARTS INSIDE

The AI playground.

Play, explore GPUs, and see how AI makes a reply.

PLAY YOUR GPU ADVENTURE QUESTS + ARCADE

Small token.
Infinite
possibilities.

Help a GPU do the math. Find room for an AI model. Follow a message as it becomes a reply. Six quests, an arcade game, and 3D hardware to explore.

Play at your pace·Keyboard or touch·No background needed
PLAY TO LEARN
MEET YOUR WORLD · GET HANDS-ON
THE HARDWARE STUDIO

GB300 NVL72Blackwell Ultra

START OUTSIDE. GO ALL THE WAY IN.

A whole system. Select an assembly, then step inside.

CAMERA VIEWS
Model viewpoints, not photographs
Preparing your hardware…
INTERACTIVE LEARNING MODEL not exact CAD
Drag 360° · pinch / scroll to zoomASSEMBLED
A GUIDED LOOK INSIDE

Meet the machine in five stops.

Open the assembly, find its working memory, and follow the compute hierarchy. Start when you are ready.

100%
18 compute trays · 9 switch trays. Inventory is not allocation. Positions and flow timing are illustrative.
CONNECTED WITH NVLINK72 GPUs

An entire team of GPUs. Software chooses how to share the work.

YOUR QUEST LINE

Don’t just read it. Make it work.

0 / 6 finished this visit
Fictional Token Quest adventure world: a token crosses bridges between parallel-work, matrix, memory and cache islands.YOUR CAMPAIGN WORLD · ILLUSTRATIVE GAME ART
Try something and watch what changes. Every quest has a hint and an example to help you.
Token Dash, the arcade.

45 seconds of play, with three pauses for questions.

  1. Catch the dataCollect what the GPU needs
  2. Avoid obstaclesKeep your streak alive
  3. Answer quick questionsLearn as you play
CURIOSITY HAS NO WRONG TURNS

Choose your next adventure.

One universe. Your way in.
PLAY45 SEC

Chase the next token.

Catch data, avoid obstacles, and answer three questions. See if you can beat your best score.

EXPLORE15 MACHINES

Meet the machines.

Turn the 3D model, separate its parts, and zoom into a chip. Links let you compare it with NVIDIA's own diagrams.

FOLLOW A REQUESTGO DEEPER

How does AI make a reply?

Send a message and follow it step by step. Watch it become numbers the model can work with, then a reply you can read.

Every great run starts with one little hop.

No account. No homework. Just a little curiosity.

KEEP LEARNING

Want to learn more?

Step-by-step GPU guides from Kubesimplify.

8 guides
START WITH THE SERIES

7 Days of Local LLM

Run your first AI model on your own computer. Then learn about the hardware and software that make it work.

5 published parts of a planned 7 · Verified Sep 9, 2026
Explore the series
Field guidePart 1

Day 1: The Local LLM Revolution. Why Your Desk Just Became the New Datacenter

What do you need to run AI on your own computer? Start with the model, machine and software. See what a small DGX Spark can and cannot do.

DGX Spark and software for running AI models locally.

13 min read
Field guidePart 2

Day 2: Anatomy of an LLM Inference Request. From Prompt to Answer, Step by Step

Follow a message from the moment you send it to the reply on your screen. See why starting a reply and continuing it involve different work.

A decoder-model request explained with DGX Spark examples.

26 min read
Field guidePart 3

Day 3: The DGX Spark Unpacked. GB10, Unified Memory, sm_121, and the One Reason This Hardware Exists

Look inside DGX Spark. Find out what GB10 is, how the CPU and GPU share memory, and why fitting a big model does not always mean a fast reply.

DGX Spark's GB10 platform; not a B200 data-center GPU.

19 min read
Field guidePart 4

Day 4: Quantization Demystified. BF16, FP8, NVFP4, MXFP4, INT4, GGUF, and Why It All Matters

What do the letters on a model download mean? Learn how storing numbers with fewer bits can save space, and why you still need to check answer quality and speed.

Numerical formats and model storage, with DGX Spark examples.

28 min read
Field guidePart 5

Day 5: Local LLM Inference Engines, Wrappers, and What to Pick

Meet the software that runs local models, including Ollama, llama.cpp and vLLM. Compare how they work and what they suit. No one tool is best for every task.

The author's DGX Spark setup. Results can change with the software version and model format.

56 min read
BenchmarkMeasured results

Bonsai 27B on RTX PRO 6000 vs DGX Spark: what actually works

See tests of compressed Bonsai models on Spark and RTX PRO 6000 Server Edition. Find out which setups worked, which failed, and how the authors measured speed.

RTX PRO 6000 Blackwell Server Edition, not this game's Workstation Edition. Results describe the linked test, not our simulator.

14 min read
RunbookMulti-GPU serving

Running a big LLM across multiple GPUs with vLLM

Run a model using several GPUs with vLLM. Follow the memory checks, choose how to split the work, and learn from problems the authors met on their RTX PRO 6000 server.

Four allocated RTX PRO 6000 Blackwell Server Edition GPUs in an eight-GPU host. This is not an NVL72 system or the catalog's Workstation Edition.

39 min read
ReferencePlain-language reference

The Local LLM Glossary: Every Term, Flag, and Number in Plain English

Stuck on a word or a setting? Look it up here, then return to the guide. Covers the terms used when running and testing local AI models.

Local LLM vocabulary and serving examples, not a hardware specification sheet.

19 min read

These guides share the authors' experience, not official NVIDIA specifications. Test results apply to the setup in each article. We have not repeated those tests here. Reading times come from the publisher. Links open in a new tab.