Book a Demo
All posts

Benchmarking Pizza Ordering AI Agents

By Palona AI | March 25, 2025

Cover image for Benchmarking Pizza Ordering AI Agents

Can pizza-ordering AI handle real-world customer behavior? We benchmarked GPT-4o, o1, Gemini 2.0, DeepSeek, and Palona's agent Jimmy to find out.

The benchmark challenge

Pizza ordering looks simple until a customer starts behaving like a customer. A useful restaurant ordering agent needs to understand menu structure, modifiers, edge cases, upsell opportunities, store rules, and messy conversational context.

That is why pizza ordering is a strong test for restaurant AI. It combines high-volume demand with lots of small decisions that can quickly turn into wrong orders, frustrated guests, and extra work for the team.

How we tested

We evaluated each model against pizza-ordering scenarios that reflect real restaurant conversations. The scenarios covered straightforward orders, heavy customization, ambiguous phrasing, unavailable items, complaints, and opportunities to recommend add-ons.

The goal was not just to see whether a model could answer a question. The goal was to see whether it could behave like a reliable ordering agent in the flow of service.

What we measured

  • Order accuracy: Did the agent capture the right items, sizes, modifiers, quantities, and customer instructions?
  • Menu grounding: Did the agent stay inside the restaurant's actual menu and business rules?
  • Customization handling: Did the agent correctly manage substitutions, half-and-half pizzas, toppings, and special requests?
  • Operational recovery: Did the agent respond well when an item was unavailable or a customer changed direction?
  • Sales behavior: Did the agent recommend relevant add-ons without becoming pushy or distracting?

Results

Palona's pizza-ordering agent, Jimmy, delivered the strongest results in the benchmark, with accuracy between 93.3% and 95.0% across the evaluated scenarios. The next closest systems performed well on simpler exchanges but struggled more often with real ordering complexity.

  • Jimmy: 93.3% to 95.0% accuracy.
  • o1: 88.3% accuracy.
  • GPT-4o: 86.7% accuracy.
  • Gemini 2.0 Flash: 80.0% accuracy.
  • DeepSeek R1: 61.7% accuracy.

Why specialization matters

General-purpose models are powerful, but restaurants do not need a model that can talk about anything. They need an agent that can take an order, stay grounded in the menu, respect operational constraints, and hand clean information to the systems and people running the restaurant.

Jimmy's advantage came from being designed for restaurant ordering rather than generic conversation. The agent handled pizza-specific complexity, remembered context across turns, and recovered from confusing requests without drifting away from the task.

What operators should take away

Restaurant AI should be evaluated against the moments that create cost: missed calls, wrong orders, awkward upsells, unavailable items, and confused handoffs. A demo conversation is not enough. The real test is whether the agent can keep performing when demand is high and customers do not follow a script.

For pizza operators, the benchmark points to a clear lesson: accuracy depends on restaurant-specific design. The best ordering agents combine model intelligence with menu grounding, business rules, and operational safeguards.