Core Event: Multi-Agent LLM System Deployment

DoorDash has deployed a production-grade multi-agent large language model (LLM) system to automatically clean up expired feature flags in its codebase. The system has completed validation—no commercial release or third-party access is involved, as this is purely an internal engineering efficiency tool.
Key facts:
- Orchestrator Agent: Powered by Claude Sonnet, manages task coordination and metadata fetching
- Cleanup Agent: Powered by Claude Opus, executes code modifications and verification
- Execution context: Isolated Git Worktrees; up to 4 agents per repository
- Timeout limit: 1 hour per cleanup task
- Evaluation scope: 50 sampled expired flags (not full-scale production rollout)
Architecture and Cleanup Logic

DoorDash’s experiment platform manages over 60,000 feature flags across approximately 623 repositories, with around 2,300 new flags added monthly. The team identified over 1,000 expired flags—defined as those without modifications in 90 days, still referenced in code, not archived retired, and not explicitly excluded.
The workflow consists of two stages:
- Orchestration: The orchestrator agent uses Google’s Agent Development Kit and queries the experiment platform via Model Context Protocol (MCP) to retrieve flag metadata (e.g., rollout percentage, targeting values), generates reports for engineer approval
- Cleanup: The cleanup agent locates all references (including test files), determines removal strategy, modifies code, executes builds and tests, performs JaCoCo patch coverage and Detekt static analysis, and creates a Pull Request only after passing all checks
Expired flags identified daily trigger automated Jira ticket creation, enabling human review before code changes proceed.
Evaluation Results
Evaluation results for 50 sampled expired flags:
| Metric | Value | Notes |
|---|---|---|
| Successful PR generation rate | 90% (45/50) | 5 required further human intervention |
| Average cleanup duration | 13.8 minutes | Labor estimate: 1–2 hours |
| Cost per cleanup | $4.79 | Estimated from sampled tasks |
| Immediate merge (first PR) | 31 cases | |
| Revised merge (subsequent PR) | 14 cases | |
| Simple flag success rate | 100% | |
| Medium-complexity flag success rate | 94% | |
| High-complexity flag success rate | 85% |
Counterintuitive finding: While Uber’s open-source Piranha uses AST-based rule transformations, DoorDash found it fails for its dependency-injection-based wrapper pattern—the relationship between flags and business logic is semantic, not syntactic. This highlights the fundamental distinction between rule-based and LLM-driven approaches.
All five human-intervention cases involved deep call chains and cross-interface parameter passing. No bugs or regressions were introduced across the failure cases.
Implementation Guidance

- Adopt if: Your feature flag inventory exceeds 1,000, you operate a mature experiment platform, and have end-to-end automation (build, test, coverage, static analysis)
- Wait if: Your codebase exhibits heavy dependency injection with tight cross-module coupling, or lacks standardized MCP access capability, as hallucination risk and debugging overhead may outweigh gains
Notably, DoorDash runs Gradle with daemon disabled to prevent state leakage across isolated Worktrees—a pragmatic detail worth replicating for similar systems.
Final Notes
DoorDash’s work demonstrates that LLMs can achieve high accuracy with zero regressions in code refactoring when constrained by engineering safeguards. Its approach—bounding agent behavior with explicit timeouts and parallelism limits—offers a reusable template for industrial LLM deployment. The project has been accepted at ICSME 2026 Industry Track.
