πŸ€₯ PINOCCHIO BENCH

do budget AI models actually do what they say? β€” a live fabrication benchmark

This exists because a free model once told me "Very good, sir β€” it's scheduled" about a calendar event that never existed. Each model below runs real assistant tasks with mock tools. We record what it called versus what it claimed. Claiming success without calling anything is a fabrication β€” the Pinocchio.
Loading results…