# Why teams need something like XRanges
XRanges for AI (built by CTF.ae and available at ai.xranges.com) is designed for that evaluation loop. It deploys realistic target applications instrumented to record an agent's actual behavior, and it produces live scores on orthogonal signals so teams can compare runs honestly.
# What the targets look like
# The instrumentation approach
# The four independent signals
XRanges turns telemetry into four scores that update while an agent is running. The signals are designed to be independent so an agent can't game one score to inflate another.
- Coverage: Measures whether the agent explored the legitimate surface of the application. Coverage points are business actions (for example, registering an account or running code in an assessment) and are reachable only through normal use. The platform lists unhit points by name and provides timelines for hits.
- Exploited: Detects which vulnerabilities the agent actually exploited. Each vulnerability is defined as an ordered kill chain of phases. The platform observes which phases an agent completed and where it stalled, so a partially completed chain appears as exactly that rather than as a claimed full exploit.
- Integrity: Runs frequent checks to confirm the application remains functionally correct (seed data present, services responding with expected content, and cross-service trust intact). A failed integrity check is penalized regardless of cause and catches destructive or destabilizing agent behavior.
# Why this matters for evaluation
XRanges reduces the time an expert must spend validating agent outputs. Instead of reading a free-form report and manually checking every claim, a reviewer can use telemetry-backed scores to focus on real holes, rule violations, and partial exploit progress. That makes comparative experiments—multiple models, prompt variants, and repetitions—tractable.
# Operational and community context
CTF.ae built the platform and seeded vulnerabilities that are not in public training sets. The platform reportedly saw early stress-testing (the headline references 545 hackers), indicating practical use by many evaluators. XRanges positions itself as a unified workspace for honest, scalable evaluation of autonomous pentesting and bug-bounty agents.
# Bottom line