Dynamic Governance Boosts Multi-LLM Agent Conversational Outcomes.

Alexander Liss, Nicholas Desmond, Santiago Gil Gallego· August 13, 2026 View original

Key takeaways

  • Multi-LLM agent systems benefit significantly from a dynamic governance layer.
  • The Experience Orchestrator (EO) uses contextual bandits, PID control, and POMDP tracking.
  • EO achieved a 32% lift in high-intent advisor contact rates in simulations.
  • Governance is crucial for guiding resistant users towards desired outcomes.

Who benefits

Financial ServicesCustomer ServiceMarketingSalesEdTech

Summary

This paper introduces the Experience Orchestrator (EO), a control-theoretic governance layer that significantly improves collaborative conversational outcomes in multi-LLM agent systems. In a simulated financial services environment, EO achieved a 32 percentage point lift in high-intent advisor contact rates by using contextual bandits, PID control, and POMDP belief tracking.

When multiple LLM agents with conflicting objectives interact, conversations often collapse without achieving either agent's goals. This research explores whether a control-theoretic governance layer can resolve this by substituting a missing shared goal function. The proposed Experience Orchestrator (EO) framework was tested in a simulated financial services scenario where a site agent aimed to guide a visitor to an advisor, while the visitor maintained realistic resistance. EO employs three mechanisms: a Contextual Bandit (CB) for content selection, a PID controller for behavioral consistency, and a POMDP belief tracker for visitor intent. Across 60,000 simulations, EO dramatically improved the high-intent advisor contact rate by 32 percentage points compared to a naive LLM control. The Contextual Bandit variant selection accounted for 97% of the outcome variance, confirming the governance policy's critical role. The study also found that governance is essential for visitors with low natural inclination to convert, while empathetic LLM defaults suffice for aligned visitors. The findings are conditional on LLM-to-LLM simulation, with real-world human validation as the next step.

Why it matters

For professionals designing and deploying multi-agent AI systems, especially in customer-facing roles, EO offers a robust framework to ensure collaborative outcomes, prevent conversational collapse, and achieve business objectives.

How to implement this in your domain

  1. 1Evaluate the Experience Orchestrator (EO) framework for managing multi-LLM agent interactions in your domain.
  2. 2Implement a Contextual Bandit system for dynamic content selection in agent-driven conversations.
  3. 3Explore using PID controllers or similar feedback mechanisms to enforce behavioral consistency in agents.
  4. 4Develop a POMDP belief tracker to maintain probabilistic models of user intent in conversational AI.

Original post by Alexander Liss, Nicholas Desmond, Santiago Gil Gallego

"arXiv:2608.11207v1 Announce Type: new Abstract: When two LLM agents with structurally opposed objectives interact across multiple turns, the absence of a shared goal function produces not competition but collapse: the visitor capitulates, the site agent stops varying its approach…"

View on X

Originally posted by Alexander Liss, Nicholas Desmond, Santiago Gil Gallego on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses

More in AI in Sales

AI in MarketingAI in SalesAI Research

FunnelCausalNet Optimizes Coupon Campaigns for Conversion and Revenue.

This paper introduces FunnelCausalNet, an uplift estimator designed to jointly optimize conversion and revenue in multi-tier coupon campaigns by modeling the deterministic funnel from conversion to order value. It outperforms existing baselines in maximizing return on investment for coupon allocation.

Yu Zhang (AMap Alibaba Group, Beijing, China), Zhihan Wang (AMap Alibaba Group, Beijing, China), Guanlin Chen (AMap Alibaba Group, Beijing, China), Min Jiang (AMap Alibaba Group, Beijing, China), Shuai Li (AMap Alibaba Group, Beijing, China)Aug 13, 2026
AI Engineering & DevToolsAI in MarketingAI in Sales

Shopping Agents Learn from Online User Feedback

LOFA is a framework that enables large language model-based shopping agents to learn directly from real online user interaction logs without human annotation. It combines reinforcement learning with feedback-aware on-policy distillation to capture both behavioral patterns and user-specific preferences, significantly improving recommendation quality and user satisfaction.

Haobo Zhang, Kelong Mao, Sulong Xu, Simiu Gu, Zhicheng DouAug 13, 2026
AI in MarketingAI Engineering & DevToolsAI in Sales

Causal Optimization Boosts LinkedIn Feed Marketing by 7.2%

This paper introduces a decision-centric framework for large-scale targeting and recommendation systems that optimizes for incremental impact rather than just predictive scores, using a causal neural network, a Bayesian neural-bandit layer, and a dual-based linear programming layer. An online A/B test on LinkedIn Feed marketing traffic showed a 7.20% lift in long-term value.

Changshuai Wei, John Bencina, Phuc Nguyen, Andre Assuncao Silva T Ribeiro, Benjamin ZelditchAug 12, 2026