Blog

Notes, write-ups, and project updates from the den.

Why I Built BenchPup: Making LLM Benchmarking Human-Readable

BenchPup started as a frustration with opaque benchmark scores. This post explains the problem it solves, the design philosophy behind it, and why human-readable metrics matter for local AI workflows.

Read More →