← All projects

ML Systems · Analytical and measured tooling

TensorForge

A two-part toolkit: explicit accelerator performance models in Core, then measured workload and regression evidence in TensorForge Ops.

Why I built it

TensorForge began with a question: can an accelerator estimate remain inspectable all the way from workload shape to memory traffic? TensorForge Ops added the boundary the equations cannot cross—what a real PyTorch workload did on actual hardware.

Architecture

Core decomposes GEMM, Transformer, and Conv2D workloads; maps work onto PE arrays; checks SRAM feasibility; compares tile-residency schedules; calculates DRAM traffic; and applies roofline reasoning. A deterministic bounded search ranks only supplied candidates. Ops records PyTorch latency, throughput, memory, telemetry, calibration, regression, sizing, and impact artifacts.

Python · PyTorch · roofline analysis · SRAM tiling · regression policy

Where the model stops

Different residency schedules produce different traffic because they preserve different operands across loop nests. Those are predictions under explicit assumptions, not measurements. Ops checks workload and artifact identity before comparison, applies metric-specific regression policy, and keeps analytical explanation distinct from measured evidence.

What I actually tested

  • hand-derived GEMM, traffic, and roofline cases
  • Transformer and Conv2D decomposition consistency
  • deterministic design-space ranking and feasibility rejection
  • benchmark artifact identity and regression-policy outcomes
  • impact reports combining measurement, sizing, telemetry, and limitations

Boundary and trade-off

Core is not cycle-accurate and does not claim model-accuracy validation. Real measurements remain hardware- and software-specific; analytical results explain a hypothesis rather than replacing benchmark evidence.

Related writing