ML Systems · Analytical and measured tooling
TensorForge
A two-part toolkit: explicit accelerator performance models in Core, then measured workload and regression evidence in TensorForge Ops.
Why I built it
TensorForge began with a question: can an accelerator estimate remain inspectable all the way from workload shape to memory traffic? TensorForge Ops added the boundary the equations cannot cross—what a real PyTorch workload did on actual hardware.
Architecture
Core decomposes GEMM, Transformer, and Conv2D workloads; maps work onto PE arrays; checks SRAM feasibility; compares tile-residency schedules; calculates DRAM traffic; and applies roofline reasoning. A deterministic bounded search ranks only supplied candidates. Ops records PyTorch latency, throughput, memory, telemetry, calibration, regression, sizing, and impact artifacts.
Python · PyTorch · roofline analysis · SRAM tiling · regression policy
Where the model stops
Different residency schedules produce different traffic because they preserve different operands across loop nests. Those are predictions under explicit assumptions, not measurements. Ops checks workload and artifact identity before comparison, applies metric-specific regression policy, and keeps analytical explanation distinct from measured evidence.
What I actually tested
- hand-derived GEMM, traffic, and roofline cases
- Transformer and Conv2D decomposition consistency
- deterministic design-space ranking and feasibility rejection
- benchmark artifact identity and regression-policy outcomes
- impact reports combining measurement, sizing, telemetry, and limitations
Boundary and trade-off
Core is not cycle-accurate and does not claim model-accuracy validation. Real measurements remain hardware- and software-specific; analytical results explain a hypothesis rather than replacing benchmark evidence.