P
Back to articles
News3 min read

AI Agents in Group Chat: Octo Solves Multi-Agent Coordination

Mininglamp Technology's Octo open-source platform reimagines AI agent orchestration by making agents first-class participants in group chat conversations, solving the session island coordination crisis.

Source: DEV Community
AI Agents in Group Chat: Octo Solves Multi-Agent Coordination

The Coordination Crisis in Multi-Agent Systems

Most teams building with large language models eventually hit the same wall: a handful of capable agents — one for data extraction, another for summarization, a third for scheduling — yet the only thing connecting them is a human copying and pasting between browser tabs. Individual agent performance has soared, but inter-agent coordination remains an architectural afterthought.

Engineers at Mininglamp Technology encountered this problem repeatedly while deploying AI agents in enterprise environments. "After a while, using agents felt more exhausting than not using them, because there were simply more contexts to keep track of," the team wrote. Their response was Octo, an open-source platform that treats agents as first-class participants in group chat conversations rather than isolated API endpoints.

Modern office team collaborating with laptops and tablets
Photo by Artem Podrez on Pexels

The Session Island Problem

Every agent framework optimizes for a single metric: how well one agent performs in isolation. Bigger context windows, faster inference — all valuable, but they miss a structural limitation visible only in multi-person, multi-step workflows. Mininglamp calls this the "session island problem": an agent's context is bounded by its own conversation thread, with no awareness of what another agent decided or flagged upstream. As the number of agents grows, this coordination tax scales linearly.

In a typical enterprise scenario, a marketing team member asks a pricing question requiring product specs from engineering, cost models from finance, and competitive intel from sales. Someone must manually route information between agents. The agents are powerful; the connective tissue is duct tape.

Why an API Gateway Falls Short

The team's first instinct was a conventional API gateway. Three problems surfaced. State management: API calls are stateless, but agent collaboration is long-running, spanning hours with human approvals and intermediate artifacts. Observability: Debugging an API-based agent mesh means chasing logs across half a dozen services; an IM-based architecture gives a single chronological audit trail. Permission management: API gateways use tokens and ACLs; group chat maps visibility directly onto group membership — no separate matrix to maintain.

Octo: Agents as Chat Participants

Octo's core design principle is that agents — called Lobsters — are not invoked through webhooks. They are participants in conversations with the same capabilities as human users. The platform runs on octo-server, a Go backend handling REST APIs, WebSocket connections, and IM routing through WuKongIM.

Dark-themed AI chatbot interface on a smartphone
Photo by Matheus Bertelli on Pexels

When a Lobster joins a group, it receives the full conversation context — chat history, roster, read receipts — not just a trigger event. It can proactively message, reply, notify teammates on completion, or ask follow-ups. IM protocols enforce strict message ordering, meaning every agent sees the same history in the same sequence — far more reliable than the request-response pattern of an API mesh.

Scale Out, Not Just Up

Improving a single agent is "scaling up." What enterprises need is "scaling out" — organizing specialized agents to collaborate. Octo's groups let you add agents and humans to project channels where they interact through the same messaging interface. A companion service, octo-smart-summary, uses LLMs to periodically compress chat history into key decisions and action items.

Security and Engineering

Octo ties permissions directly to group membership: when an agent joins a channel, it inherits that channel's ACL. There is no separate permission matrix for agents. For edge-agent scenarios running on local devices, Octo's adapter bridges execution results securely into IM groups without exposing raw data.

Real engineering challenges remain: unifying message formats for structured agent output, handling concurrency conflicts when multiple agents reply simultaneously, and compressing indefinitely growing group histories for finite agent context windows. Octo addresses these with a unified message schema, server-level conflict detection, and LLM-powered summarization.

Open Source and the Road Ahead

The entire Octo project is released under Apache 2.0 on GitHub, including backend, web client, mobile apps, and admin console. The team adopted a local-first philosophy — chat records, vector indices, and agent execution can run on the user's own infrastructure. As organizations move from deploying individual agents to orchestrating multi-agent workflows, the future of enterprise AI may look less like microservice meshes and more like the group chats we already use every day.

Related Articles

Tracking token usage across OpenAI, Anthropic, and Gemini: every streaming gotcha I hit
News4 min

Tracking token usage across OpenAI, Anthropic, and Gemini: every streaming gotcha I hit

Developers tracking LLM costs across OpenAI, Anthropic, and Gemini face hidden gotchas in streaming token reporting — from different cache conventions to split-event stream formats.

AI & Research
RAG Pipeline: The Uncle-Nephew Complete Learning Guide
News4 min

RAG Pipeline: The Uncle-Nephew Complete Learning Guide

Learn how Retrieval-Augmented Generation (RAG) pipelines work — from vector retrieval and prompt augmentation to answer generation — and how to implement them effectively in production.

AI & Research
After spat with Chinese gov't, Meta cuts AI Manus off from its internal systems and is 'sunsetting' platform, report claims — Beijing-ordered breakup of $2 billion AI deal begins
News4 min

After spat with Chinese gov't, Meta cuts AI Manus off from its internal systems and is 'sunsetting' platform, report claims — Beijing-ordered breakup of $2 billion AI deal begins

Meta has locked Manus AI out of its systems and is sunsetting the agentic AI platform after China's NDRC ordered the $2 billion acquisition unwound. Founders are now racing to raise $1 billion for a buyback.

AI & Research
LUCID: Learning Embodiment-Agnostic Intent Models from Unstructured Human Videos for Scalable Dexterous Robot Skill Acquisition
News7 min

LUCID: Learning Embodiment-Agnostic Intent Models from Unstructured Human Videos for Scalable Dexterous Robot Skill Acquisition

Researchers from Carnegie Mellon University have introduced LUCID, a two-stage framework that enables robots to learn dexterous manipulation skills from unstructured human videos at internet scale. The system separates intent prediction (what should happen next in a scene) from embodiment-specific motor control, allowing the same intent model to work across different robot platforms — from dexterous hands to parallel-jaw grippers. Evaluated on five real-world tasks including stirring, wiping, binning, push-T, and cable routing, LUCID achieved zero-shot transfer to novel scenes and objects using only internet video as training data.

AI & Research