Skip to content

Accuracy = Efficiency × Resources

The most efficient, portable runtime for neural networks.

Research keeps climbing a ladder of abstraction: training your own model, then fine-tuning a downloaded one, then prompting an API you never open. Each step buys productivity, but unlike a programming language or an operating system, these abstractions leak — and there is still research that needs the whole stack, not just the top of it.

Forge exists on the premise that understanding this technology comes from building it. One Rust crate, one set of WGSL kernels, no CUDA toolchain and no Python interpreter between you and the model trains and runs on every GPU wgpu reaches, from a workstation to a browser tab to the edge devices you own end to end. Scroll down and it will run on yours, and tell you how fast.

verified — CPU and WGPU logits agree to 8.4e-5, and GPT-2's output matches HuggingFace transformers token for token (tests/gpt2_e2e.rs, on an RTX A5000).

Watch it think

What it chose

The model's own probability for each candidate, before temperature and top-k reshape the distribution it samples from — so the character it picks is not always the one at the top.

Where it looked

Attention for the position just computed, over the 32 characters before it. Bar heights are relative to the strongest position in that row; hover one for the exact probability.

block
head

Nothing has been computed yet

Every number here is produced by shakespeare-char, a 10.77M parameter model trained from scratch with Forge. Start it and the numbers arrive one character at a time — 6.7 MB of q4-quantized weights, fetched once and then cached by your browser.

The whole matrix, every position

Rows are queries, columns are keys, and the triangle is the causal mask — no position may attend to the future. Brightness is p0.45 so small weights stay visible; the number you hover is the raw probability.

No server does the work — your GPU does. Forge needs no SharedArrayBuffer, because rayon is native-only, so this page is plain static files.

What a token costs

10.77 million parameters is not a small language model by the industry's use of the phrase — it is three orders of magnitude below it. The research question is what that floor actually buys, and the only answer worth printing is the one measured on the machine in front of you.

Decode speed
tok/s
one point per token · dashed line, the run's average
Your GPU
Prompt encode
tok/s

Timed over the prompt of your next run.

Weights on disk
MB
fp32
on disk

Measured against the same tensors at four bytes each.

Parameters
M
embeddings
attention
MLP

Nothing here is filled in until you press Run above — every figure is produced by your own hardware, so none of it can go stale.

Runs where you are

One set of WGSL kernels, four places to run them. The browser tab and the workstation take the same code path — not an equivalent one, the same one.

Where Backend What it needs
Linux workstation wgpu → Vulkan a Vulkan ICD (mesa or the vendor driver)
macOS, Apple silicon wgpu → Metal nothing beyond the OS
Windows wgpu → D3D12 nothing beyond the OS
This page wasm32 → WebGPU Chrome/Edge 113+ or Safari 26+

The model, and how to run it

One Rust crate, small enough to read in an afternoon — all of it is on GitHub. Below is the model the page above just ran, and the two commands that serve this page on your own machine.

The model: shakespeare-char

Trained from scratch with Forge — one GPU, one command, no pretrained weights and no Python in the loop. 10.77M parameters: 6 blocks × 6 heads × 384 dims, a 65-character vocabulary, 256 characters of context, LM head weight-tied to the token embedding.

Held-out validation loss , against nanoGPT's published 1.4697 for the same configuration.

The longest passage it shares with the text it was trained on is 24 characters, and 98% of the words it invents are real words from the corpus. It is composing, not reciting.

Run this page locally

./scripts/build_site.sh
./scripts/serve_web.sh
# plain static files — no COOP/COEP headers needed

30 WGSL kernels in Rust

The complete compute surface, generated from shaders/ at build time so this list cannot drift from the code.

Forward — 22

  • add
  • dropout
  • embedding
  • gather_nll
  • gelu
  • gemv
  • gemv_reduce
  • kv_append
  • l2_norm
  • l2_norm_bwd
  • layernorm
  • matmul
  • mean_pool
  • mean_pool_bwd
  • merge_heads
  • scale
  • softmax
  • softmax_masked
  • split_heads
  • sum_rows
  • unmerge_heads
  • unsplit_heads

Backward — 6

  • ce_bwd
  • gelu_bwd
  • layernorm_bwd_dp
  • layernorm_bwd_dx
  • scatter_add
  • softmax_bwd

Optimizer — 2

  • adamw
  • sumsq

Train it yourself

Tiny Shakespeare, from scratch

./scripts/download_shakespeare.sh

# 6 blocks / 6 heads / 384 dims, 65-token vocab
cargo run --release --example train_shakespeare -- \
  --backend wgpu

What comes out

A safetensors checkpoint and its vocabulary. Quantize it to the .fzm q4 file this page loads with --example to_fzm, or point --model at it and the demo above runs your weights instead — the forward pass, the KV cache and the trace do not know the difference.

Quick start

Generate — no download

# the 6.7 MB char model is in the repo
cargo run --release --example generate -- \
  --model assets/shakespeare_char --prompt "ROMEO:"

Same binary, real weights: ./scripts/download_gpt2.sh, then drop the --model flag to run GPT-2 124M — the run tests/gpt2_e2e.rs checks against HuggingFace.

Browse models in the terminal

# forge-top: model browser + live dashboard
# tokens/s, VRAM, GPU util, temperature, power
cargo run --release -p forge-top -- --path models/

A crate in tools/, not a feature of the runtime — so nothing that draws a terminal can reach a dependent's build.