Back to Articles
Part 1: Simulation, Harness and Frontier LLM testing
Most AI benchmarks ask a model to answer. DukaanBench asks a model to operate.
In DukaanBench, a language model runs a small Indian kirana store for 30 simulated days. Every morning it receives the shop state, recent sales and misses, inventory, cash, trust, weather, customer signals, khata exposure, active marketing, and a fixed neighborhood profile. It must return one executable JSON action before the shop opens. The backend then simulates customers, stockouts, payments, khata, waste, marketing effects, trust movement, and reward.
Across the 30 days, the AI goes through the ordinary pressure of shopkeeping: it starts with limited cash and shelves, decides what to reorder, protects fast-moving essentials, handles perishable stock, chooses whether to discount or market products, reminds khata customers, and then watches simulated customers either get served or walk away disappointed. Each day leaves a mark. Cash, inventory, trust, customer memory, and missed demand carry into the next morning.









