# What happened Researchers at Group-IB identified a phishing platform called Balonx that includes a module named CallFlow. CallFlow uses multiple commercial AI services together to fully automate voice-phishing (vishing) calls. The system stages an entire conversational attack using synthetic speech, real-time transcription, and a large language model to generate live responses.
# How CallFlow works CallFlow chains four commercial components: GPT-4o-mini for dialogue generation, ElevenLabs's AI voice generator for a consistent speaking persona, OpenAI Voice for text-to-speech integration, and OpenAI Whisper for real-time transcription of victims' spoken replies. Group-IB describes a fabricated bank representative called Carolina (an ElevenLabs voice profile) as the voice victims hear.
When a call connects, the victim hears the synthetic agent. The victim's spoken answers are captured and transcribed by Whisper. Transcripts feed into GPT-4o-mini, which produces the next voice response. The generated text is rendered through the synthetic voice and delivered to the victim, creating a continuous conversational loop that most people perceive as a human operator.
# Operator role and campaign management Although the conversation itself can be fully automated, Balonx includes an administrative interface for human operators to monitor and manage campaigns. The CallFlow dashboard displays active calls, daily volume, success rates, calls in queue, active campaigns, total phone records, and average call duration. Operators can watch in real time and intervene if the AI encounters difficulty.
Group-IB's investigation revealed targeted campaigns aimed at several banks and a telephone number database tied to a specific bank brand, suggesting attackers maintain curated contact lists for focused attacks.
# Why this raises concern Vishing historically relied on human agents, which limited scale. By automating the interaction, CallFlow removes that constraint: a single operator can oversee hundreds of simultaneous fraudulent calls without participating in each conversation. Group-IB warns that as this architecture spreads, financial institutions could face a much larger volume of voice scams with improved linguistic quality.
The use of widely available AI services means attackers do not need custom models or deep voice-engineering expertise to produce convincing, real-time voice conversations. The platform's combination of voice synthesis and real-time transcription makes it practical to run high-volume, context-aware calls that adapt to victims' answers.
# Current scope and trajectory Group-IB observed Balonx targeting users in Mexico at the time of their research. The researchers expect the same automation model to expand internationally as criminal networks adopt and distribute the kit.
# Practical steps for organizations and individuals Security awareness training that covers voice-based social engineering can help employees recognize and resist vishing attempts. Organizations that rely on phone-based authentication or support should reassess those workflows and consider additional verification steps for high-risk transactions. Monitoring for anomalous call patterns and educating customers about legitimate contact channels may reduce success rates of automated vishing campaigns.
# Bottom line Balonx demonstrates that commercially available AI components can be combined into a turnkey vishing platform that conducts convincing, adaptive conversations at scale. Targeted sectors and banks should treat voice phishing as an elevated risk and update detection, authentication, and training strategies accordingly.