The art & science of multiprovider AI workflows

Claude, Gemini, GPT, and Grok now supervise and fact-check each other in real-world collaborative and competitive hierarchies. This workshop asks: where is the science to hone and advance that art?

๐Ÿ“… Date & venue TBD Apply to participate โ†’

A powerful new practice is emerging amidst an explosion of partially-but-imperfectly aligned AI deployments: the art of coordinating AI workflows involving multiple AI providers complementing and supervising each other. Concretely, Claude, Gemini, GPT, and Grok models are all now being used in collaborative and competitive management hierarchies, in real-world applications where they supervise and fact-check each other.

Where is the science to hone and advance this art? The time is ripe for empiricism. Indeed, Google DeepMind, Schmidt Sciences, the Cooperative AI Foundation, and ARIA have all announced funding opportunities in multi-agent AI trust and safety.

And yet, there is little consensus on what are the most important problems, solutions, questions, or answers in multi-principal AI alignment. Indeed, this lack of consensus can be seen in the opinions of leading AI models themselves: aside from MultiAgentBench, we found literally zero overlap in top-ten lists constructed by Claude, Gemini, GPT, and Grok models, when prompted for publicly-available scientific metrics tracking the capacities of AI models deployed in multi-provider collaborative workflows (multi-chat). By contrast, when prompted for general alignment metrics, 8 items from their top-10 lists were overlapping contributions (multi-chat).

The aim of this workshop is to improve the state of scientific consensus on what is most important and tractable to measure next in multiprovider AI systems. Crucially, attendees are all expected to have real-world experience building multiprovider AI workflows โ€” whether for research, engineering, business, or other productive use โ€” so our scientific discussions can remain grounded in real-world impacts already unfolding today.

If you have this sort of experience, and would like to help build scientific consensus in this area, please apply below to participate!

Application

We hope to get back to you within two weeks of your application with an update regarding your admission.

1 Contact details
2 Your experience
How many hours have you spent building such workflows (and/or using workflows you built or customized for yourself)? *
3 Questions for consensus
4 Your contribution
5 A link about you

About

This workshop is organized and co-sponsored by FAR.AI, Encultured AI, and the UC Berkeley Center for Human-Compatible AI.