Zapier expanded its AI automation platform to include new model releases from OpenAI, Anthropic, and Google, sharing updated results from its AutomationBench, a proprietary measure of how well models handle multi-step business workflows rather than isolated prompts.
Zapier's model catalog shows Gemini 3.7 Flash (High) leading AutomationBench at 30.44% task completion and $0.61 per task, ahead of Opus 5 (Max) at 26.94% and $1.27 per task. New additions include GPT-5.6 Sol ($30 per million output tokens), Terra ($15), and Luna ($6) from OpenAI, and Fable 5.0 ($50 per million) from Anthropic for long-horizon agentic work.
OpenAI's three new models target distinct workflow profiles. Sol is built for high-stakes, one-shot processes: compliance reviews, approval chains, and situations where the correct move may be to pause or escalate rather than proceed. Terra handles multi-source assembly tasks, pulling data from a CRM, email, and collaboration tools before producing a coherent output. Luna addresses high-volume, rule-based operations such as discount logic, lead tagging, and ticket routing, where cost accumulates quickly at scale.
On the Anthropic side, four Opus variants hold four of the top five AutomationBench positions. The newly listed Fable model sits above Opus in the hierarchy for the most demanding long-horizon agentic tasks, while Sonnet covers everyday professional work at a mid-tier price point and Haiku handles high-throughput, latency-sensitive pipelines at the budget end.
Google's Gemini range spans cost-efficient Flash variants suited to real-time classification and chat, up to Pro-tier models for complex reasoning and large-scale data synthesis. Multimodal input across text, images, audio, and video applies throughout the Gemini family, a practical consideration for workflows handling media alongside structured data.
Beyond the three major families, Zapier supports direct integrations with Grok for real-time web and social data, DeepSeek for cost-efficient technical pipelines, Mistral for multilingual tasks, and OpenRouter, which routes requests to dozens of providers from a single connection. Zapier's built-in AI step, AI by Zapier, bundles core models from OpenAI, Anthropic, and Google and allows teams to swap providers via dropdown without rebuilding the underlying Zap.
AutomationBench tasks are deliberately structured around real business friction: irrelevant data injected into inputs, key information placed behind required tool calls, similar naming conventions to create plausible wrong answers, and policy rules with overriding priority over simpler heuristics. The benchmark evaluates only the end state of each task, so a model scores only when the workflow completes correctly and without side effects, regardless of how many tool calls it made to get there. The methodology targets multi-step workflow completion specifically, which is why Gemini's leading position in Zapier's leaderboard differs from its placement in most general-purpose model rankings.













