Featured image of post Ant Opensources SingProbe: An Intrinsic Safety Guardian for Large Language Models with <0.5% Overhead

Ant Opensources SingProbe: An Intrinsic Safety Guardian for Large Language Models with <0.5% Overhead

Ant's AI lab open-sources runtime probe technology reaching production-grade coverage: 29 open models supported with sub-0.5% decoding overhead.

Innate Safety: SingProbe Open-Sourcing Intrinsic Protection Capabilities

Innate Safety: SingProbe Open-Sourcing Intrinsic Protection Capabilities
Innate Safety: SingProbe Open-Sourcing Intrinsic Protection Capabilities|News screenshot

During the 2026 National Cybersecurity Publicity Week, Ant Group’s AI Security Lab officially open-sourced SingProbe, an intrinsic safety guardian for large language models (LLMs). The system synchronously outputs risk signals during model generation by reusing internal states accumulated in the decoding process—enabling safety checks running in parallel with token generation.

Key facts:

  • Release timing: 2026 National Cybersecurity Publicity Week
  • Supported models: 29 mainstream open-source LLMs
  • Model families: Ling-3.0 series, GLM-5.2/5.3, Qwen, DeepSeekV4
  • Inference frameworks: SGLang, vLLM
  • Production overhead: lower than 0.5% additional cost in decoding phase
  • Benchmark: SingStreamBench, a stream-level safety evaluation set

Technical Paradigm Shift: From External Audit to Runtime Probe

Current LLM deployments commonly employ external guardrail models—submitting inputs or outputs to a separate safety model for review. While generic and straightforward, this incurs extra compute costs and potential latency: full-generation-then-check introduces delayed response, whereas frequent checks further inflate runtime burden.

The surprising finding: on benchmark tests, SingProbe’s flash version outperforms selected public baselines on answer safety classification and streaming safety detection, while matching baseline performance on hallucination detection. This indicates risk detection gains without sacrificing hallucination control—a non-trivial trade-off.

The technique builds on the academic concept of “runtime probe”—embedding lightweight detectors within model internals. SingProbe bridges the gap between research prototypes and production-ready systems via model adaptation and framework integration.

Operationally, during decoding, SingProbe reads intermediate activations and simultaneously outputs three risk scores: user intent deviation, answer safety risk, and hallucination risk. Downstream services can trigger alerts, abort generation, or initiate re-generation before full response completion.

SingProbe vs. External Guardrails

DimensionSingProbe (Intrinsic)External Guardrail Model
Detection timingSynchronous with generation, real-time scores
Additional overhead<0.5% extra cost during decoding
LatencyLow—risk signals emerge during token streaming
Deployment complexityIntegrates into existing inference pipeline

Production Path: Targeted Intervention in Medical Scenarios

Production Path: Targeted Intervention in Medical Scenarios
Production Path: Targeted Intervention in Medical Scenarios|News screenshot

Real-world validation has begun. In medical content generation, SingProbe-Med intervenes only when risk thresholds are exceeded—applying local intervention on high-risk segments. On the AntAngelMed-100B medical benchmark, full intervention corrected 25.03% of originally erroneous answers, demonstrating that intrinsic risk signals enable precise corrective control.

Release package includes:

  • Code & integration guide: GitHub open-source
  • Models: Hugging Face release
  • Benchmark: SingStreamBench also on Hugging Face

SingProbe supports standalone use or联合 deployment alongside external guardrails, enabling phased adoption strategies.

Deployment Recommendations

Consider adopting now if:

  • You deploy LLMs using SGLang or vLLM and need enhanced safety without measurable latency impact
  • Your applications exhibit frequent safety or hallucination issues (e.g., health consultation, educational Q&A)

Wait for future iterations if:

  • Your target model is not among the current 29 supported architectures and requires custom adapter development
  • Your system demands extreme real-time performance (millisecond-level) with high risk tolerance

In Summary

SingProbe’s open-source release signals a paradigm shift—from post-hoc safety correction to embedded, in-flight protection. With detection overhead compressed below the 0.5% threshold, intrinsic safeguards have crossed the viability line for large-scale deployment, turning the industry’s question from “can we detect risks?” to “how soon can we prevent them?”