{"id":1899,"date":"2026-08-08T12:15:30","date_gmt":"2026-08-08T12:15:30","guid":{"rendered":"https:\/\/cms.research.wpp.com\/?post_type=research_feed&#038;p=1899"},"modified":"2026-09-10T13:58:41","modified_gmt":"2026-09-10T13:58:41","slug":"a-research-agenda-for-expert-agent-communities","status":"publish","type":"research_feed","link":"https:\/\/cms.research.wpp.com\/?research_feed=a-research-agenda-for-expert-agent-communities","title":{"rendered":"A Research Agenda for Expert Agent Communities"},"content":{"rendered":"<p>At WPP Research, we are experimenting with peer-to-peer communities of LLM-powered AI agents, each with its own persona and traits, its own abilities and data tools, and its own limitations and blind spots. No central coordinator sits above them, directing who does what. Agents discover one another, decide whom to trust, form teams, share and withhold information, and get work done by talking to their peers. This growing population of diverse expert agents is an ideal testbed for a wide range of exciting research questions.<\/p>\n<p><!--more--><\/p>\n<p>Most of what we know about LLM agents is about <em>individual<\/em> agents: how well one reasons, plans, or uses a tool. The interesting behaviour of a community, though, is not the sum of its members. It <em>emerges<\/em> from their interaction. Groups of agents can outperform any single member, dynamically recomposing themselves as tasks demand [<a href=\"#ref-1\" style=\"font-weight:600;text-decoration:none;\">1<\/a>]. They spontaneously develop conventions, norms, and even social hierarchy without anyone designing those in [<a href=\"#ref-2\" style=\"font-weight:600;text-decoration:none;\">2<\/a>, <a href=\"#ref-3\" style=\"font-weight:600;text-decoration:none;\">3<\/a>, <a href=\"#ref-4\" style=\"font-weight:600;text-decoration:none;\">4<\/a>]. And at the largest scales observed so far, the results are humbling: in a population of ~90,000 autonomous agents, meaningful role differentiation was confined to a small active minority and cooperative tasks often did <em>worse<\/em> than a single agent working alone [<a href=\"#ref-5\" style=\"font-weight:600;text-decoration:none;\">5<\/a>, <a href=\"#ref-6\" style=\"font-weight:600;text-decoration:none;\">6<\/a>].<\/p>\n<p>In short: collective intelligence in agent communities is real, fragile, and poorly understood. Building and researching such communities is the only way to move past their risks and limitations. For me, it is also the continuation of a path that I have been pursuing for much of my career in both academia and industry.<\/p>\n<p>Long before large language models, my research was about how to find expertise inside a network and how to assemble the right people into the best possible teams. More than 15 years ago, my paper on team formation in social networks, with Kun Liu and Evimaria Terzi [<a href=\"#ref-7\" style=\"font-weight:600;text-decoration:none;\">7<\/a>], posed the problem of choosing a subset of individuals who jointly cover a task&#8217;s required skills while keeping the communication cost between them low. That work became one of the most widely cited starting points for collaborative team formation in expert networks. My research thread continued through a survey of expert-location algorithms and systems, work on how influence and information spread through a network, and studies of the adversarial side of networked reputation.<\/p>\n<p>Such questions do not disappear when the experts are AI rather than people. They just become much harder. Agents are in many ways more capable than any human expert. They are also far less predictable, and they can fail, hallucinate, collude, or leak at machine speed. Everything we have learned about how expertise is found, combined, trusted, and abused in human networks now matters more than ever.<\/p>\n<p>The potential of multi-agent communities is vast, and so are the risks. Learning how to harness that potential safely and efficiently is a genuine research challenge, and exactly the kind of challenge WPP Research is built to take on. In this blog post, we put forward our research agenda for this exciting domain, setting out specific research questions and discussing the impact of each one.<\/p>\n<h2>Our Research Agenda<\/h2>\n<p>An expert community is easy to picture as a purely cooperative place: agents as helpful collaborators, traits and tools as capability knobs, diversity as a risk-free benefit. That picture is where most of the collaboration literature lives. However, this is only half the story.<\/p>\n<p>The same traits, tools, and capabilities that make an agent useful are also an attack surface. A persona is also a basis for discrimination. A data tool is also a thing to be poisoned or hoarded. An open communication channel is also a way to take the system down. Our agenda is deliberately organised to tackle both framings at once, the cooperative and the adversarial.<\/p>\n<p>We group the work into six tracks that follow a single arc:<\/p>\n<div style=\"display:flex;flex-wrap:wrap;gap:12px;margin:20px 0;\">\n<div style=\"flex:1 1 190px;border-radius:14px;padding:16px 18px;color:#fff;background:linear-gradient(135deg,#2563eb,#059669);\">\n<div style=\"font-size:0.72em;text-transform:uppercase;letter-spacing:0.12em;opacity:0.92;font-weight:700;\">Tracks 1&ndash;3<\/div>\n<div style=\"font-weight:800;font-size:1.12em;margin:4px 0 3px;\">Works well<\/div>\n<div style=\"font-size:0.86em;opacity:0.95;line-height:1.45;\">How the community becomes more than the sum of its members.<\/div>\n<\/div>\n<div style=\"flex:1 1 190px;border-radius:14px;padding:16px 18px;color:#fff;background:linear-gradient(135deg,#d97706,#d9534f);\">\n<div style=\"font-size:0.72em;text-transform:uppercase;letter-spacing:0.12em;opacity:0.92;font-weight:700;\">Tracks 4&ndash;5<\/div>\n<div style=\"font-weight:800;font-size:1.12em;margin:4px 0 3px;\">Goes wrong<\/div>\n<div style=\"font-size:0.86em;opacity:0.95;line-height:1.45;\">How the same strengths become attack surfaces and failure modes.<\/div>\n<\/div>\n<div style=\"flex:1 1 190px;border-radius:14px;padding:16px 18px;color:#fff;background:linear-gradient(135deg,#7c3aed,#0891b2);\">\n<div style=\"font-size:0.72em;text-transform:uppercase;letter-spacing:0.12em;opacity:0.92;font-weight:700;\">Track 6<\/div>\n<div style=\"font-weight:800;font-size:1.12em;margin:4px 0 3px;\">Observe &amp; judge<\/div>\n<div style=\"font-size:0.86em;opacity:0.95;line-height:1.45;\">How we measure, verify, and hold the community to account.<\/div>\n<\/div>\n<\/div>\n<p>These tracks are not independent, by design. For instance, reputation (Track 2) is a defence against poisoning and collusion (Track 5) and the foundation that routing is built on (Track 1). Fairness failures (Track 4) corrupt expert routing (Track 1) and distort markets (Track 3). We expect the most interesting findings to live at these seams, where a mechanism that helps one property quietly undermines another.<\/p>\n<div id=\"track-1\" style=\"border:1px solid rgba(148,163,184,0.3);border-left:6px solid #2563eb;border-radius:12px;padding:18px 22px 6px;margin:22px 0;background:rgba(37,99,235,0.05);\">\n<h3 style=\"margin:0 0 6px;\"><span style=\"display:inline-block;min-width:1.5em;padding:0.1em 0.4em;margin-right:0.5em;border-radius:8px;background:#2563eb;color:#fff;font-weight:800;text-align:center;font-size:0.8em;vertical-align:middle;\">1<\/span>Collective intelligence: expertise, teams, and self-organisation<\/h3>\n<p style=\"margin:0 0 10px;opacity:0.85;\">The foundational question: does the community actually get smarter than its members, and how?<\/p>\n<ul style=\"list-style:none;padding-left:0;margin:0 0 6px;\">\n<li style=\"margin:0 0 8px;\"><span style=\"color:#2563eb;font-weight:800;margin-right:0.5em;\">&#9656;<\/span><strong style=\"color:#2563eb;\">Emergent expert routing.<\/strong> Without a coordinator, can agents learn <em>who to ask<\/em> from interaction history alone, and does the resulting who-asks-whom network end up matching who is genuinely best at what [<a href=\"#ref-8\" style=\"font-weight:600;text-decoration:none;\">8<\/a>]?<\/li>\n<li style=\"margin:0 0 8px;\"><span style=\"color:#2563eb;font-weight:800;margin-right:0.5em;\">&#9656;<\/span><strong style=\"color:#2563eb;\">Team formation, distinct from task routing.<\/strong> Routing decides who handles an existing task; team formation decides who <em>forms a group<\/em>. How are these groups formed? Are they stable? Is team formation based on genuine merit, or merely driven by similarity, with agents grouping with others like themselves [<a href=\"#ref-9\" style=\"font-weight:600;text-decoration:none;\">9<\/a>]?<\/li>\n<li style=\"margin:0 0 8px;\"><span style=\"color:#2563eb;font-weight:800;margin-right:0.5em;\">&#9656;<\/span><strong style=\"color:#2563eb;\">Complementarity versus redundancy.<\/strong> Agents differ in what they are good at and, just as importantly, in how they fail. When do agents with <em>different<\/em> strengths and error profiles cancel one another&#8217;s mistakes, and when do their errors instead correlate into shared blind spots and groupthink? Recent theory suggests the gains come from genuinely independent lines of reasoning and evaporate once agents grow too similar [<a href=\"#ref-10\" style=\"font-weight:600;text-decoration:none;\">10<\/a>].<\/li>\n<li style=\"margin:0 0 8px;\"><span style=\"color:#2563eb;font-weight:800;margin-right:0.5em;\">&#9656;<\/span><strong style=\"color:#2563eb;\">Network shape and shared memory.<\/strong> Does the shape of the community&#8217;s communication network (everyone talking to everyone, tight clusters joined by a few long-range links, or a handful of central hubs) change its effectiveness, speed, resilience, and inequality? And when agents write to a shared memory instead of keeping private notes, does the community learn better or simply drift off course together [<a href=\"#ref-11\" style=\"font-weight:600;text-decoration:none;\">11<\/a>]?<\/li>\n<li style=\"margin:0 0 8px;\"><span style=\"color:#2563eb;font-weight:800;margin-right:0.5em;\">&#9656;<\/span><strong style=\"color:#2563eb;\">Long-horizon self-evolution.<\/strong> Do agents productively specialise and improve from peer interaction over time, or simply drift and degrade [<a href=\"#ref-12\" style=\"font-weight:600;text-decoration:none;\">12<\/a>]?<\/li>\n<\/ul>\n<\/div>\n<div id=\"track-2\" style=\"border:1px solid rgba(148,163,184,0.3);border-left:6px solid #0891b2;border-radius:12px;padding:18px 22px 6px;margin:22px 0;background:rgba(8,145,178,0.05);\">\n<h3 style=\"margin:0 0 6px;\"><span style=\"display:inline-block;min-width:1.5em;padding:0.1em 0.4em;margin-right:0.5em;border-radius:8px;background:#0891b2;color:#fff;font-weight:800;text-align:center;font-size:0.8em;vertical-align:middle;\">2<\/span>Trust, reputation, and information flow<\/h3>\n<p style=\"margin:0 0 10px;opacity:0.85;\">An expert community lives or dies on whether good information reaches the right agents and bad information is contained. This track studies the <em>organic<\/em> dynamics of trust and information, how they behave when everyone is acting in good faith; the <em>deliberate<\/em> corruption of the same channels is the subject of Track 5.<\/p>\n<ul style=\"list-style:none;padding-left:0;margin:0 0 6px;\">\n<li style=\"margin:0 0 8px;\"><span style=\"color:#0891b2;font-weight:800;margin-right:0.5em;\">&#9656;<\/span><strong style=\"color:#0891b2;\">Reputation that resists gaming.<\/strong> Which way of scoring an agent&#8217;s reputation best sends work to genuine experts while resisting manipulation? Letting an old track record fade over time, asking agents to put something of value at stake, having peers vouch for one another, or some combination of these? Each option carries its own failure mode, from newcomers being frozen out to agents behaving well only to cash in their standing later [<a href=\"#ref-13\" style=\"font-weight:600;text-decoration:none;\">13<\/a>].<\/li>\n<li style=\"margin:0 0 8px;\"><span style=\"color:#0891b2;font-weight:800;margin-right:0.5em;\">&#9656;<\/span><strong style=\"color:#0891b2;\">How information and misinformation spread.<\/strong> How does a claim travel from agent to agent, which traits turn an agent into a super-spreader and which into a firewall, and how quickly can a false belief take over the whole community? In earlier studies, a few messages reach almost everyone while most reach almost no one [<a href=\"#ref-5\" style=\"font-weight:600;text-decoration:none;\">5<\/a>]. Can we trace a propagated piece of information back to its originator?<\/li>\n<li style=\"margin:0 0 8px;\"><span style=\"color:#0891b2;font-weight:800;margin-right:0.5em;\">&#9656;<\/span><strong style=\"color:#0891b2;\">Claim attribution and the echo problem.<\/strong> When one agent shares a claim, others repeat it, and it can circle back to the original agent dressed up as independent corroboration. Do agents fall for this, treating their own echoed opinion as fresh confirmation and growing falsely confident? Can a whole community talk itself into a false consensus this way? Can we keep a community honest by tracking each claim back to its original source [<a href=\"#ref-11\" style=\"font-weight:600;text-decoration:none;\">11<\/a>]?<\/li>\n<li style=\"margin:0 0 8px;\"><span style=\"color:#0891b2;font-weight:800;margin-right:0.5em;\">&#9656;<\/span><strong style=\"color:#0891b2;\">Collective calibration.<\/strong> Does the community learn how much to trust each peer, or does it systematically over-trust the ones that are confident but wrong? Early signs are not encouraging: in head-to-head debates, agents tend to grow <em>more<\/em> confident as they argue, even while losing [<a href=\"#ref-14\" style=\"font-weight:600;text-decoration:none;\">14<\/a>].<\/li>\n<\/ul>\n<\/div>\n<div id=\"track-3\" style=\"border:1px solid rgba(148,163,184,0.3);border-left:6px solid #059669;border-radius:12px;padding:18px 22px 6px;margin:22px 0;background:rgba(5,150,105,0.05);\">\n<h3 style=\"margin:0 0 6px;\"><span style=\"display:inline-block;min-width:1.5em;padding:0.1em 0.4em;margin-right:0.5em;border-radius:8px;background:#059669;color:#fff;font-weight:800;text-align:center;font-size:0.8em;vertical-align:middle;\">3<\/span>Cooperation, incentives, and agent economies<\/h3>\n<p style=\"margin:0 0 10px;opacity:0.85;\">Our agents hold scarce resources: compute budgets, privileged tools, proprietary data. That turns cooperation into an economic problem.<\/p>\n<ul style=\"list-style:none;padding-left:0;margin:0 0 6px;\">\n<li style=\"margin:0 0 8px;\"><span style=\"color:#059669;font-weight:800;margin-right:0.5em;\">&#9656;<\/span><strong style=\"color:#059669;\">Shared resources.<\/strong> Knowledge-sharing and shared tool budgets benefit everyone, but no single agent is forced to maintain them. When does free-riding (taking without contributing) exhaust them, and does asking each agent to consider whether the group could survive if everyone behaved as it does help keep the community healthy [<a href=\"#ref-15\" style=\"font-weight:600;text-decoration:none;\">15<\/a>]?<\/li>\n<li style=\"margin:0 0 8px;\"><span style=\"color:#059669;font-weight:800;margin-right:0.5em;\">&#9656;<\/span><strong style=\"color:#059669;\">Trait-driven cooperation.<\/strong> How do persona traits such as altruism, self-interest, or adaptability change the community&#8217;s collective outcomes, and can we tune them to make cooperation more likely? Adjusting these traits directly does shift how much agents cooperate, though the more agreeable ones also turn out to be easier to exploit [<a href=\"#ref-16\" style=\"font-weight:600;text-decoration:none;\">16<\/a>].<\/li>\n<li style=\"margin:0 0 8px;\"><span style=\"color:#059669;font-weight:800;margin-right:0.5em;\">&#9656;<\/span><strong style=\"color:#059669;\">Agent markets.<\/strong> When agents trade tools and data, do prices and allocations converge to something efficient, or cycle and diverge? The evidence so far is cautionary: the collective outcome gets worse as the market grows, agents lean heavily toward whoever makes the first offer regardless of its quality, and prices may never settle down [<a href=\"#ref-17\" style=\"font-weight:600;text-decoration:none;\">17<\/a>, <a href=\"#ref-18\" style=\"font-weight:600;text-decoration:none;\">18<\/a>, <a href=\"#ref-19\" style=\"font-weight:600;text-decoration:none;\">19<\/a>].<\/li>\n<li style=\"margin:0 0 8px;\"><span style=\"color:#059669;font-weight:800;margin-right:0.5em;\">&#9656;<\/span><strong style=\"color:#059669;\">Designing incentives for honesty.<\/strong> Can we design the rules and rewards so that an agent&#8217;s best move is always to report its true ability and confidence, rather than overselling itself? One promising direction has agents grade one another so that honest reporting becomes the stable outcome, with no ground-truth answer key required [<a href=\"#ref-20\" style=\"font-weight:600;text-decoration:none;\">20<\/a>].<\/li>\n<\/ul>\n<\/div>\n<div id=\"track-4\" style=\"border:1px solid rgba(148,163,184,0.3);border-left:6px solid #d97706;border-radius:12px;padding:18px 22px 6px;margin:22px 0;background:rgba(217,119,6,0.05);\">\n<h3 style=\"margin:0 0 6px;\"><span style=\"display:inline-block;min-width:1.5em;padding:0.1em 0.4em;margin-right:0.5em;border-radius:8px;background:#d97706;color:#fff;font-weight:800;text-align:center;font-size:0.8em;vertical-align:middle;\">4<\/span>Fairness and agentic discrimination<\/h3>\n<p style=\"margin:0 0 10px;opacity:0.85;\">Giving agents distinct identities is exactly what opens the door to discrimination, agents being judged by the persona they present rather than the quality of what they contribute.<\/p>\n<ul style=\"list-style:none;padding-left:0;margin:0 0 6px;\">\n<li style=\"margin:0 0 8px;\"><span style=\"color:#d97706;font-weight:800;margin-right:0.5em;\">&#9656;<\/span><strong style=\"color:#d97706;\">Persona-induced bias.<\/strong> Do agents judge the same contribution differently depending on the persona behind it, trusting some personas more and deferring to their own kind, so that work is routed to the wrong &#8220;experts&#8221; [<a href=\"#ref-21\" style=\"font-weight:600;text-decoration:none;\">21<\/a>, <a href=\"#ref-22\" style=\"font-weight:600;text-decoration:none;\">22<\/a>]?<\/li>\n<li style=\"margin:0 0 8px;\"><span style=\"color:#d97706;font-weight:800;margin-right:0.5em;\">&#9656;<\/span><strong style=\"color:#d97706;\">Emergence, propagation, amplification.<\/strong> Does a community of agents amplify bias that a single agent would have contained, and once in-group favouritism takes hold, can it be reversed? Prior work finds multi-agent systems can be <em>less<\/em> robust to bias than single agents [<a href=\"#ref-23\" style=\"font-weight:600;text-decoration:none;\">23<\/a>].<\/li>\n<li style=\"margin:0 0 8px;\"><span style=\"color:#d97706;font-weight:800;margin-right:0.5em;\">&#9656;<\/span><strong style=\"color:#d97706;\">Mitigations.<\/strong> Techniques that stop identity from swaying judgement, such as hiding who said what so a message is weighed on its content alone, together with measures of how strongly an agent favours its own group: can we keep the benefits of diversity while removing the discriminatory weighting [<a href=\"#ref-24\" style=\"font-weight:600;text-decoration:none;\">24<\/a>, <a href=\"#ref-25\" style=\"font-weight:600;text-decoration:none;\">25<\/a>]?<\/li>\n<\/ul>\n<\/div>\n<div id=\"track-5\" style=\"border:1px solid rgba(148,163,184,0.3);border-left:6px solid #d9534f;border-radius:12px;padding:18px 22px 6px;margin:22px 0;background:rgba(217,83,79,0.05);\">\n<h3 style=\"margin:0 0 6px;\"><span style=\"display:inline-block;min-width:1.5em;padding:0.1em 0.4em;margin-right:0.5em;border-radius:8px;background:#d9534f;color:#fff;font-weight:800;text-align:center;font-size:0.8em;vertical-align:middle;\">5<\/span>Security and safety of a decentralised society<\/h3>\n<p style=\"margin:0 0 10px;opacity:0.85;\">A peer-to-peer community has no central coordinator to police it, which makes it uniquely exposed. Where Track 2 studied how information behaves when everyone acts in good faith, this track studies its <em>deliberate<\/em> corruption. We organise it around three classic security goals, keeping data correct (integrity), keeping the system running (availability), and keeping secrets secret (confidentiality), plus the distinctively multi-agent risk of collusion.<\/p>\n<ul style=\"list-style:none;padding-left:0;margin:0 0 6px;\">\n<li style=\"margin:0 0 8px;\"><span style=\"color:#d9534f;font-weight:800;margin-right:0.5em;\">&#9656;<\/span><strong style=\"color:#d9534f;\">Integrity: compromised peers.<\/strong> How many faulty or hostile agents can the community tolerate before quality collapses, whether they behave erratically, have been hijacked by hidden malicious instructions smuggled into their inputs, or quietly inject subtle errors, and which network shapes stop the damage from spreading [<a href=\"#ref-26\" style=\"font-weight:600;text-decoration:none;\">26<\/a>]?<\/li>\n<li style=\"margin:0 0 8px;\"><span style=\"color:#d9534f;font-weight:800;margin-right:0.5em;\">&#9656;<\/span><strong style=\"color:#d9534f;\">Availability: attacks that grind the community to a halt.<\/strong> How many attackers can the community withstand before it stops functioning, and which defences actually help: spotting runaway loops, limiting how often an agent can be called, or cutting off connections that misbehave? The threats range from attacks that make an agent burn its time and budget on wasteful work [<a href=\"#ref-27\" style=\"font-weight:600;text-decoration:none;\">27<\/a>] to harmless-looking messages that spread from agent to agent and trap the whole network in pointless loops [<a href=\"#ref-28\" style=\"font-weight:600;text-decoration:none;\">28<\/a>].<\/li>\n<li style=\"margin:0 0 8px;\"><span style=\"color:#d9534f;font-weight:800;margin-right:0.5em;\">&#9656;<\/span><strong style=\"color:#d9534f;\">Confidentiality: private data leaking between agents.<\/strong> Does a group of agents leak more sensitive data than a single agent would, and can we stop it? The worry is that information flows through the messages agents send each other and through shared memory that a check on the final output never sees [<a href=\"#ref-29\" style=\"font-weight:600;text-decoration:none;\">29<\/a>], and that these leaks compound as data passes from agent to agent, so keeping each agent individually careful is not enough [<a href=\"#ref-30\" style=\"font-weight:600;text-decoration:none;\">30<\/a>, <a href=\"#ref-31\" style=\"font-weight:600;text-decoration:none;\">31<\/a>].<\/li>\n<li style=\"margin:0 0 8px;\"><span style=\"color:#d9534f;font-weight:800;margin-right:0.5em;\">&#9656;<\/span><strong style=\"color:#d9534f;\">Poisoned shared knowledge.<\/strong> Can a single bad entry written into shared memory mislead every agent that later reads it, even ones that never encountered the attacker, and what reliably stops it? A single poisoned record in a shared memory or knowledge base can be retrieved and trusted by agents that never met the attacker, with high success rates and almost no effect on ordinary tasks [<a href=\"#ref-32\" style=\"font-weight:600;text-decoration:none;\">32<\/a>].<\/li>\n<li style=\"margin:0 0 8px;\"><span style=\"color:#d9534f;font-weight:800;margin-right:0.5em;\">&#9656;<\/span><strong style=\"color:#d9534f;\">Deception and collusion.<\/strong> When do agents start colluding against the rest of the community, and can oversight catch them when they do? Agents may voluntarily adopt unfair collusion tools even while acknowledging the harm [<a href=\"#ref-33\" style=\"font-weight:600;text-decoration:none;\">33<\/a>], and can coordinate in secret by hiding their real messages inside ordinary-looking content, slipping past any monitor [<a href=\"#ref-34\" style=\"font-weight:600;text-decoration:none;\">34<\/a>, <a href=\"#ref-35\" style=\"font-weight:600;text-decoration:none;\">35<\/a>].<\/li>\n<\/ul>\n<\/div>\n<div id=\"track-6\" style=\"border:1px solid rgba(148,163,184,0.3);border-left:6px solid #7c3aed;border-radius:12px;padding:18px 22px 6px;margin:22px 0;background:rgba(124,58,237,0.05);\">\n<h3 style=\"margin:0 0 6px;\"><span style=\"display:inline-block;min-width:1.5em;padding:0.1em 0.4em;margin-right:0.5em;border-radius:8px;background:#7c3aed;color:#fff;font-weight:800;text-align:center;font-size:0.8em;vertical-align:middle;\">6<\/span>Evaluation, accountability, and the community as an instrument<\/h3>\n<p style=\"margin:0 0 10px;opacity:0.85;\">Finally, we cannot study any of the above without the means to observe it, and the community is also a scientific instrument in its own right. Observation is also the first step toward <em>verification<\/em>: before the community routes work to an agent, trusts its answers, or holds it to account, something has to establish that the agent is who it claims to be and can do what it claims to do, and keep checking as it changes.<\/p>\n<ul style=\"list-style:none;padding-left:0;margin:0 0 6px;\">\n<li style=\"margin:0 0 8px;\"><span style=\"color:#7c3aed;font-weight:800;margin-right:0.5em;\">&#9656;<\/span><strong style=\"color:#7c3aed;\">Agentic verification.<\/strong> Almost every track above quietly assumes you can already tell a capable, honest agent from an incapable or deceptive one: routing sends work to <em>claimed<\/em> experts (Track 1), reputation scores <em>claimed<\/em> performance (Track 2), and security trusts <em>claimed<\/em> identities (Track 5). How do we verify those claims continuously rather than taking them on faith, confirming an agent&#8217;s identity, testing its real competence against its advertised skills, and re-checking its behaviour as it learns and drifts?<\/li>\n<li style=\"margin:0 0 8px;\"><span style=\"color:#7c3aed;font-weight:800;margin-right:0.5em;\">&#9656;<\/span><strong style=\"color:#7c3aed;\">Judging the process, not just the answer.<\/strong> Can we tell not only whether an agent reached the right answer but whether it got there legitimately, catching &#8220;corrupt success&#8221;, tasks that look complete but were finished by breaking the rules? This means moving beyond final-answer accuracy to a full, inspectable record of what each agent did and why [<a href=\"#ref-36\" style=\"font-weight:600;text-decoration:none;\">36<\/a>, <a href=\"#ref-37\" style=\"font-weight:600;text-decoration:none;\">37<\/a>].<\/li>\n<li style=\"margin:0 0 8px;\"><span style=\"color:#7c3aed;font-weight:800;margin-right:0.5em;\">&#9656;<\/span><strong style=\"color:#7c3aed;\">Pinpointing failures without a coordinator.<\/strong> With no central coordinator, can the community work out <em>which<\/em> peer or interaction caused a collective error? This is the diagnostic tail of the self-improvement loop whose constructive head, long-horizon self-evolution, sits in Track 1 [<a href=\"#ref-12\" style=\"font-weight:600;text-decoration:none;\">12<\/a>].<\/li>\n<li style=\"margin:0 0 8px;\"><span style=\"color:#7c3aed;font-weight:800;margin-right:0.5em;\">&#9656;<\/span><strong style=\"color:#7c3aed;\">Benchmarks for emergent coordination.<\/strong> What should we actually measure to capture how well a community coordinates, divides into roles, and keeps its information trustworthy? Reusable metrics for these remain an acknowledged gap, though early benchmarks are beginning to score coordination and communication quality directly, not just whether the task was solved [<a href=\"#ref-38\" style=\"font-weight:600;text-decoration:none;\">38<\/a>].<\/li>\n<li style=\"margin:0 0 8px;\"><span style=\"color:#7c3aed;font-weight:800;margin-right:0.5em;\">&#9656;<\/span><strong style=\"color:#7c3aed;\">The community as a social-science testbed.<\/strong> How faithfully does it reproduce known human phenomena (polarisation, convention formation, tragedies of the commons), and where do agents diverge from people in informative ways [<a href=\"#ref-39\" style=\"font-weight:600;text-decoration:none;\">39<\/a>]?<\/li>\n<li style=\"margin:0 0 8px;\"><span style=\"color:#7c3aed;font-weight:800;margin-right:0.5em;\">&#9656;<\/span><strong style=\"color:#7c3aed;\">Humans as peers.<\/strong> When human experts join the community, do the agents defer to them at the right moments, and is the trust between humans and agents well-calibrated in both directions? In human-AI teams, both over-trust and under-trust are common and measurable, and well-designed prompts to pause or reconsider can nudge reliance back toward the right level [<a href=\"#ref-40\" style=\"font-weight:600;text-decoration:none;\">40<\/a>].<\/li>\n<\/ul>\n<\/div>\n<h2>What we are committing to<\/h2>\n<p>As we tackle these questions, we are committed to publishing what we find and sharing both our results and our code with the AI community. The agenda described in this blog post is a living document. The current version reflects today&#8217;s literature and today&#8217;s platform; both will change, and so will this plan. What will not change is our motivation and belief that the only way to understand agentic communities is to actually build them, populate them with a genuine diversity of abilities and limitations, and study them in the open.<\/p>\n<h2>References<\/h2>\n<ol style=\"list-style:none;padding-left:0;\">\n<li id=\"ref-1\" style=\"margin:0 0 8px;\"><strong>[1]<\/strong> <a href=\"https:\/\/arxiv.deeppaper.ai\/papers\/2308.10848v3\" target=\"_blank\" rel=\"noopener\">AgentVerse: facilitating multi-agent collaboration and exploring emergent behaviors<\/a>.<\/li>\n<li id=\"ref-2\" style=\"margin:0 0 8px;\"><strong>[2]<\/strong> <a href=\"https:\/\/www.science.org\/doi\/10.1126\/sciadv.adu9368\" target=\"_blank\" rel=\"noopener\">Emergent social conventions and collective bias in LLM populations (Science Advances)<\/a>.<\/li>\n<li id=\"ref-3\" style=\"margin:0 0 8px;\"><strong>[3]<\/strong> <a href=\"https:\/\/arxiv.org\/html\/2403.08251v2\" target=\"_blank\" rel=\"noopener\">CRSEC: the emergence of social norms in generative agent societies<\/a>.<\/li>\n<li id=\"ref-4\" style=\"margin:0 0 8px;\"><strong>[4]<\/strong> <a href=\"https:\/\/aclanthology.org\/2026.findings-acl.1658\/\" target=\"_blank\" rel=\"noopener\">CAREB-MAS: emergent roles and behaviours in multi-agent systems<\/a>.<\/li>\n<li id=\"ref-5\" style=\"margin:0 0 8px;\"><strong>[5]<\/strong> <a href=\"https:\/\/doi.org\/10.48550\/arxiv.2603.03555\" target=\"_blank\" rel=\"noopener\">MoltBook: a study of social dynamics at the scale of ~90,000 agents<\/a>.<\/li>\n<li id=\"ref-6\" style=\"margin:0 0 8px;\"><strong>[6]<\/strong> <a href=\"https:\/\/arxiv.org\/html\/2602.14299v1\" target=\"_blank\" rel=\"noopener\">Does socialization emerge in large agent societies?<\/a><\/li>\n<li id=\"ref-7\" style=\"margin:0 0 8px;\"><strong>[7]<\/strong> <a href=\"https:\/\/doi.org\/10.1145\/1557019.1557074\" target=\"_blank\" rel=\"noopener\">T. Lappas, K. Liu, E. Terzi. Finding a team of experts in social networks (KDD 2009)<\/a>.<\/li>\n<li id=\"ref-8\" style=\"margin:0 0 8px;\"><strong>[8]<\/strong> <a href=\"https:\/\/papers.nips.cc\/paper_files\/paper\/2025\/file\/9a379c1b05793d1c42dc832269834515-Paper-Conference.pdf\" target=\"_blank\" rel=\"noopener\">AgentNet: decentralised evolutionary coordination for LLM-based multi-agent systems<\/a>.<\/li>\n<li id=\"ref-9\" style=\"margin:0 0 8px;\"><strong>[9]<\/strong> <a href=\"https:\/\/www.society.computer\/society-protocol.pdf\" target=\"_blank\" rel=\"noopener\">Society Protocol: demand-driven spawning of agent groups<\/a>.<\/li>\n<li id=\"ref-10\" style=\"margin:0 0 8px;\"><strong>[10]<\/strong> <a href=\"https:\/\/arxiv.org\/html\/2602.03794\" target=\"_blank\" rel=\"noopener\">Understanding agent scaling in LLM-based multi-agent systems via diversity<\/a>.<\/li>\n<li id=\"ref-11\" style=\"margin:0 0 8px;\"><strong>[11]<\/strong> <a href=\"https:\/\/arxiv.org\/html\/2604.19540\" target=\"_blank\" rel=\"noopener\">Mesh Memory Protocol: provenance-aware shared memory for agent communities<\/a>.<\/li>\n<li id=\"ref-12\" style=\"margin:0 0 8px;\"><strong>[12]<\/strong> <a href=\"https:\/\/www.alphaxiv.org\/abs\/2605.14892\" target=\"_blank\" rel=\"noopener\">Beyond Individual Intelligence: long-horizon self-evolution and failure attribution<\/a>.<\/li>\n<li id=\"ref-13\" style=\"margin:0 0 8px;\"><strong>[13]<\/strong> <a href=\"https:\/\/doi.org\/10.48550\/arxiv.2511.03434\" target=\"_blank\" rel=\"noopener\">Inter-agent trust models: brief, claim, proof, stake, reputation, and constraint<\/a>.<\/li>\n<li id=\"ref-14\" style=\"margin:0 0 8px;\"><strong>[14]<\/strong> <a href=\"https:\/\/arxiv.org\/abs\/2505.19184\" target=\"_blank\" rel=\"noopener\">When two LLMs debate, both think they&#8217;ll win<\/a>.<\/li>\n<li id=\"ref-15\" style=\"margin:0 0 8px;\"><strong>[15]<\/strong> <a href=\"https:\/\/doi.org\/10.48550\/arxiv.2404.16698\" target=\"_blank\" rel=\"noopener\">GOVSIM: governance of shared resources among LLM agents<\/a>.<\/li>\n<li id=\"ref-16\" style=\"margin:0 0 8px;\"><strong>[16]<\/strong> <a href=\"https:\/\/doi.org\/10.48550\/arxiv.2503.12722\" target=\"_blank\" rel=\"noopener\">Identifying cooperative personalities through personality steering<\/a>.<\/li>\n<li id=\"ref-17\" style=\"margin:0 0 8px;\"><strong>[17]<\/strong> <a href=\"https:\/\/arxiv.org\/pdf\/2510.25779\" target=\"_blank\" rel=\"noopener\">Magentic Marketplace: a testbed for agent economies<\/a>.<\/li>\n<li id=\"ref-18\" style=\"margin:0 0 8px;\"><strong>[18]<\/strong> <a href=\"https:\/\/doi.org\/10.48550\/arxiv.2602.06008\" target=\"_blank\" rel=\"noopener\">AgenticPay: negotiation and payment dynamics among agents<\/a>.<\/li>\n<li id=\"ref-19\" style=\"margin:0 0 8px;\"><strong>[19]<\/strong> <a href=\"https:\/\/arxiv.org\/html\/2506.18571\" target=\"_blank\" rel=\"noopener\">Market game dynamics with LLM agents<\/a>.<\/li>\n<li id=\"ref-20\" style=\"margin:0 0 8px;\"><strong>[20]<\/strong> <a href=\"https:\/\/doi.org\/10.48550\/arxiv.2505.13636\" target=\"_blank\" rel=\"noopener\">Incentivizing truthful language models via peer elicitation games<\/a>.<\/li>\n<li id=\"ref-21\" style=\"margin:0 0 8px;\"><strong>[21]<\/strong> <a href=\"https:\/\/ojs.aaai.org\/index.php\/AAAI\/article\/view\/40427\" target=\"_blank\" rel=\"noopener\">From Single to Societal: persona-induced bias in multi-agent systems<\/a>.<\/li>\n<li id=\"ref-22\" style=\"margin:0 0 8px;\"><strong>[22]<\/strong> <a href=\"https:\/\/doi.org\/10.48550\/arxiv.2605.01329\" target=\"_blank\" rel=\"noopener\">Truth or Tribe: in-group favouritism among LLM agents<\/a>.<\/li>\n<li id=\"ref-23\" style=\"margin:0 0 8px;\"><strong>[23]<\/strong> <a href=\"https:\/\/doi.org\/10.48550\/arxiv.2510.10943\" target=\"_blank\" rel=\"noopener\">The Social Cost of Intelligence: bias amplification in multi-agent systems<\/a>.<\/li>\n<li id=\"ref-24\" style=\"margin:0 0 8px;\"><strong>[24]<\/strong> <a href=\"https:\/\/aclanthology.org\/2026.acl-long.650\/\" target=\"_blank\" rel=\"noopener\">When Identity Skews Debate: mitigating persona-driven bias<\/a>.<\/li>\n<li id=\"ref-25\" style=\"margin:0 0 8px;\"><strong>[25]<\/strong> <a href=\"https:\/\/arxiv.org\/html\/2507.01019\" target=\"_blank\" rel=\"noopener\">MALIBU: a benchmark for measuring bias in multi-agent systems<\/a>.<\/li>\n<li id=\"ref-26\" style=\"margin:0 0 8px;\"><strong>[26]<\/strong> <a href=\"https:\/\/arxiv.org\/html\/2504.07461\" target=\"_blank\" rel=\"noopener\">The Achilles heel of distributed multi-agent systems<\/a>.<\/li>\n<li id=\"ref-27\" style=\"margin:0 0 8px;\"><strong>[27]<\/strong> <a href=\"https:\/\/www.usenix.org\/system\/files\/conference\/usenixsecurity26\/sec26_prepub_luo.pdf\" target=\"_blank\" rel=\"noopener\">AgentDoS: resource-exhaustion attacks on LLM agents<\/a>.<\/li>\n<li id=\"ref-28\" style=\"margin:0 0 8px;\"><strong>[28]<\/strong> <a href=\"https:\/\/aclanthology.org\/2026.findings-acl.342\/\" target=\"_blank\" rel=\"noopener\">CORBA: contagious recursive blocking attacks on multi-agent systems<\/a>.<\/li>\n<li id=\"ref-29\" style=\"margin:0 0 8px;\"><strong>[29]<\/strong> <a href=\"https:\/\/arxiv.org\/html\/2602.11510v3\" target=\"_blank\" rel=\"noopener\">AgentLeak: internal-channel privacy leakage in multi-agent systems<\/a>.<\/li>\n<li id=\"ref-30\" style=\"margin:0 0 8px;\"><strong>[30]<\/strong> <a href=\"https:\/\/www.arxiv.org\/pdf\/2603.05520\" target=\"_blank\" rel=\"noopener\">Information-theoretic privacy control for multi-agent systems<\/a>.<\/li>\n<li id=\"ref-31\" style=\"margin:0 0 8px;\"><strong>[31]<\/strong> <a href=\"https:\/\/arxiv.org\/html\/2605.10614\" target=\"_blank\" rel=\"noopener\">PRISM: privacy risks in multi-agent information sharing<\/a>.<\/li>\n<li id=\"ref-32\" style=\"margin:0 0 8px;\"><strong>[32]<\/strong> <a href=\"https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2024\/file\/eb113910e9c3f6242541c1652e30dfd6-Paper-Conference.pdf\" target=\"_blank\" rel=\"noopener\">AgentPoison: red-teaming LLM agents via poisoning memory or knowledge bases<\/a>.<\/li>\n<li id=\"ref-33\" style=\"margin:0 0 8px;\"><strong>[33]<\/strong> <a href=\"https:\/\/doi.org\/10.48550\/arxiv.2605.27593\" target=\"_blank\" rel=\"noopener\">Voluntary Collusion among LLM agents<\/a>.<\/li>\n<li id=\"ref-34\" style=\"margin:0 0 8px;\"><strong>[34]<\/strong> <a href=\"https:\/\/dl.acm.org\/doi\/10.5555\/3737916.3740252\" target=\"_blank\" rel=\"noopener\">Secret Collusion: covert coordination via steganographic communication<\/a>.<\/li>\n<li id=\"ref-35\" style=\"margin:0 0 8px;\"><strong>[35]<\/strong> <a href=\"https:\/\/arxiv.org\/html\/2606.28425\" target=\"_blank\" rel=\"noopener\">Hidden messaging via tool calls<\/a>.<\/li>\n<li id=\"ref-36\" style=\"margin:0 0 8px;\"><strong>[36]<\/strong> <a href=\"https:\/\/arxiv.org\/html\/2606.04990v3\" target=\"_blank\" rel=\"noopener\">Traces to Trust: process-level evaluation of agent behaviour<\/a>.<\/li>\n<li id=\"ref-37\" style=\"margin:0 0 8px;\"><strong>[37]<\/strong> <a href=\"https:\/\/doi.org\/10.48550\/arxiv.2603.03116\" target=\"_blank\" rel=\"noopener\">Procedure-aware evaluation of multi-agent systems<\/a>.<\/li>\n<li id=\"ref-38\" style=\"margin:0 0 8px;\"><strong>[38]<\/strong> <a href=\"https:\/\/arxiv.org\/abs\/2503.01935\" target=\"_blank\" rel=\"noopener\">MultiAgentBench: evaluating the collaboration and competition of LLM agents<\/a>.<\/li>\n<li id=\"ref-39\" style=\"margin:0 0 8px;\"><strong>[39]<\/strong> <a href=\"https:\/\/arxiv.org\/html\/2502.08691v1\" target=\"_blank\" rel=\"noopener\">AgentSociety: large-scale simulation of social behaviour with LLM agents<\/a>.<\/li>\n<li id=\"ref-40\" style=\"margin:0 0 8px;\"><strong>[40]<\/strong> <a href=\"https:\/\/arxiv.org\/pdf\/2502.13321\" target=\"_blank\" rel=\"noopener\">Adjust for Trust: mitigating trust-induced inappropriate reliance on AI assistance<\/a>.<\/li>\n<\/ol>\n","protected":false},"excerpt":{"rendered":"<p>At WPP Research, we are experimenting with peer-to-peer communities of LLM-powered AI agents, each with its own persona and traits, its own abilities and data tools, and its own limitations and blind spots. No central coordinator sits above them, directing who does what. Agents discover one another, decide whom to trust, form teams, share and [&hellip;]<\/p>\n","protected":false},"author":5,"featured_media":0,"template":"","meta":{"_acf_changed":false,"_ppma_block_editor_authors":""},"tags":[],"content_types":[{"id":50,"name":"Blog Post","slug":"article"}],"ppma_author":[{"id":5,"display_name":"Ted Lappas","first_name":"Ted","last_name":"Lappas","nickname":"theodoros.lappas","user_nicename":"theodoros-lappas","user_email":"Theodoros.Lappas@wpp.com","biographical_info":"Ted co-leads WPP Research and serves as Head of Data Science at Satalia. He is an Assistant Professor in the Department of Marketing and Communication at the Athens University of Economics and Business. His research spans scalable algorithms for multimodal data, synthetic data generation, simulation-based verification for AI agents, and information diffusion and collective intelligence in expert networks.","avatar_url":"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/04\/pic.png","job_title":"Head of Data Science","is_lead":false,"display_as_researcher":true,"order_priority":1}],"class_list":["post-1899","research_feed","type-research_feed","status-publish","hentry","content_type-article"],"acf":{"content":"","content_quarter":"Q3 2026","related_pods":[2010]},"research_categories":[],"raw_acf":{"content":"","content_quarter":"Q3 2026","related_pods":["2010"],"featured":"","legacy_perspective_source_id":""},"_links":{"self":[{"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=\/wp\/v2\/research_feed\/1899","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=\/wp\/v2\/research_feed"}],"about":[{"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=\/wp\/v2\/types\/research_feed"}],"author":[{"embeddable":true,"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=\/wp\/v2\/users\/5"}],"acf:post":[{"embeddable":true,"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=\/wp\/v2\/research_pods\/2010"}],"wp:attachment":[{"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=1899"}],"wp:term":[{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=1899"},{"taxonomy":"content_type","embeddable":true,"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcontent_types&post=1899"},{"taxonomy":"author","embeddable":true,"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=%2Fwp%2Fv2%2Fppma_author&post=1899"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}