Featured image of post DoorDash Deploys Multi-Agent LLM System for Automated Expired Feature Flag Cleanup

DoorDash Deploys Multi-Agent LLM System for Automated Expired Feature Flag Cleanup

DoorDash uses multi-agent LLM system to auto-clean expired feature flags, averaging 13.8 minutes and $4.79 per cleanup.

Core Event: Multi-Agent LLM System Deployment

Core Event: Multi-Agent LLM System Deployment
Core Event: Multi-Agent LLM System Deployment|News screenshot

DoorDash has deployed a production-grade multi-agent large language model (LLM) system to automatically clean up expired feature flags in its codebase. The system has completed validation—no commercial release or third-party access is involved, as this is purely an internal engineering efficiency tool.

Key facts:

  • Orchestrator Agent: Powered by Claude Sonnet, manages task coordination and metadata fetching
  • Cleanup Agent: Powered by Claude Opus, executes code modifications and verification
  • Execution context: Isolated Git Worktrees; up to 4 agents per repository
  • Timeout limit: 1 hour per cleanup task
  • Evaluation scope: 50 sampled expired flags (not full-scale production rollout)

Architecture and Cleanup Logic

Architecture and Cleanup Logic
Architecture and Cleanup Logic|News screenshot

DoorDash’s experiment platform manages over 60,000 feature flags across approximately 623 repositories, with around 2,300 new flags added monthly. The team identified over 1,000 expired flags—defined as those without modifications in 90 days, still referenced in code, not archived retired, and not explicitly excluded.

The workflow consists of two stages:

  1. Orchestration: The orchestrator agent uses Google’s Agent Development Kit and queries the experiment platform via Model Context Protocol (MCP) to retrieve flag metadata (e.g., rollout percentage, targeting values), generates reports for engineer approval
  2. Cleanup: The cleanup agent locates all references (including test files), determines removal strategy, modifies code, executes builds and tests, performs JaCoCo patch coverage and Detekt static analysis, and creates a Pull Request only after passing all checks

Expired flags identified daily trigger automated Jira ticket creation, enabling human review before code changes proceed.

Evaluation Results

Evaluation results for 50 sampled expired flags:

MetricValueNotes
Successful PR generation rate90% (45/50)5 required further human intervention
Average cleanup duration13.8 minutesLabor estimate: 1–2 hours
Cost per cleanup$4.79Estimated from sampled tasks
Immediate merge (first PR)31 cases
Revised merge (subsequent PR)14 cases
Simple flag success rate100%
Medium-complexity flag success rate94%
High-complexity flag success rate85%

Counterintuitive finding: While Uber’s open-source Piranha uses AST-based rule transformations, DoorDash found it fails for its dependency-injection-based wrapper pattern—the relationship between flags and business logic is semantic, not syntactic. This highlights the fundamental distinction between rule-based and LLM-driven approaches.

All five human-intervention cases involved deep call chains and cross-interface parameter passing. No bugs or regressions were introduced across the failure cases.

Implementation Guidance

Implementation Guidance
Implementation Guidance|News screenshot

  • Adopt if: Your feature flag inventory exceeds 1,000, you operate a mature experiment platform, and have end-to-end automation (build, test, coverage, static analysis)
  • Wait if: Your codebase exhibits heavy dependency injection with tight cross-module coupling, or lacks standardized MCP access capability, as hallucination risk and debugging overhead may outweigh gains

Notably, DoorDash runs Gradle with daemon disabled to prevent state leakage across isolated Worktrees—a pragmatic detail worth replicating for similar systems.

Final Notes

DoorDash’s work demonstrates that LLMs can achieve high accuracy with zero regressions in code refactoring when constrained by engineering safeguards. Its approach—bounding agent behavior with explicit timeouts and parallelism limits—offers a reusable template for industrial LLM deployment. The project has been accepted at ICSME 2026 Industry Track.