← Back to Writing

When the Performance Model Meets the GPU

An analytical model can calculate exactly the answer implied by its assumptions. That does not make the answer a measurement.

TensorForge began by making accelerator reasoning explicit for GEMM, Transformer, and Conv2D workloads: operation counts, arithmetic intensity, processing-element geometry, SRAM capacity, tile schedules, DRAM traffic, and roofline bounds. TensorForge Ops later put those predictions beside measured workloads and regression evidence.

What the analytical model actually answers

TensorForge Core takes explicit workload and hardware descriptions. For GEMM it derives operations and traffic, checks whether candidate tiles fit SRAM, maps work onto a PE array, and compares c-resident, a-resident, and b-resident schedules. Transformer and Conv2D decompose into the same inspectable building blocks.

The bounded search ranks only supplied candidates under the model. It does not discover arbitrary hardware, simulate instructions, or infer compiler behavior.

Why an exact equation can still be incomplete

The Core model does not hide its boundaries. A roofline estimate selects a compute or bandwidth bound from supplied peak rates. A tiling schedule calculates traffic under a stated residency and loop-order model. Design-space exploration ranks only the bounded candidates the user supplied.

Those calculations explain why one candidate wins inside the model. They do not simulate every cache, instruction, kernel-launch, compiler, synchronization, or contention effect on a real GPU.

That separation matters. Adding an unexplained correction factor could make one dataset look closer while making the model harder to reason about. I would rather preserve a traceable lower-level explanation and label what it omits.

Identity before comparison

A measured regression result is meaningful only when the artifacts describe comparable work. TensorForge Ops records workload identity and treats mismatches as errors rather than performance changes. It distinguishes a policy failure from an operational error, so missing or incompatible evidence cannot masquerade as a slow candidate.

The regression layer can compare latency, throughput, and peak allocated memory under explicit per-metric policy. The impact report then combines that measured decision with analytical context and sizing feasibility without pretending they are the same kind of evidence.

model: why should this mapping behave this way?
measurement: what happened on this hardware and software?
policy: is that change acceptable?

Calibration is not permission to overfit

TensorForge Ops records measurements with workload and device context, then compares later runs through regression gates. Calibration can make a model more useful on observed hardware, but it can also hide a bad abstraction if I tune every constant to one device. I keep the analytical output and measured evidence separate so a disagreement remains visible. The project predicts performance structure; it does not claim cycle accuracy or validate model quality. A gate is evidence about the recorded workload, device, software environment, and threshold—not every deployment.

A real comparison

I calibrated the empirical roofline with measured compute and memory probes on one RTX 3050 Laptop GPU, then compared holdout workloads against measured PyTorch p50 latency. The small 128×128×128 GEMM was the clearest disagreement: the calibrated prediction was 552 ns, while the recorded p50 was 48.18 µs. The larger square GEMM was closer—185.18 µs predicted versus 245.97 µs measured—but still not exact.

I did not tune a correction coefficient until the small case looked good. The model has no kernel-launch, framework, detailed cache, or synchronization term, so fixed overhead is a plausible explanation for why it dominates the smaller operation. The experiment does not prove that explanation, and the validation document labels it accordingly.

model expected: roofline lower bound from measured device ceilings
measurement showed: small GEMM much farther from that bound than large GEMM
question: which omitted costs grow important as useful work shrinks?

That was the moment measurement took authority. The model remained useful as an inspectable lower-bound argument, but the GPU run decided what happened on that device. Five workloads on one laptop are not evidence of general predictive accuracy.

Disagreement is information

If the model predicts a memory-bound improvement and the benchmark does not show it, the benchmark is not wrong for failing to validate the story. The disagreement points to an omitted effect, a workload mismatch, insufficient measurement quality, or a bad assumption.

The right response is to inspect those boundaries, not promote the analytical estimate into measured fact. TensorForge's validation artifacts and limitations travel with results for that reason.

Where the model stops

Core does not model instruction issue, detailed cache behavior, synchronization, kernel launch overhead, compiler selection, or contention. It is not cycle-accurate, and the repository does not claim general prediction accuracy.

What precision means to me now

I once wanted a performance model to produce the answer. I now see its more durable role as producing an inspectable argument.

Real measurements decide what happened on a particular system. Regression policy decides whether that evidence blocks a change. The analytical model supplies a hypothesis and helps explain the result. Keeping those responsibilities separate makes each claim narrower, but far more useful.

Related project: TensorForge case study →