# What OTel 2.0 is and why AT&T deployed it
# Model design and training data AT&T built OTel 2.0 on Google Gemma 4 31B-IT (a 31-billion-parameter foundation). The project processed over one trillion tokens during preparation and narrowed the training set to about 400 billion tokens focused on telecom engineering and standards. Sources include 3GPP, ETSI, ITU, O-RAN, CAMARA, and TM Forum. That dataset choice increases domain relevance and reduces reliance on general web or personal-data sources.
# Hardware and software stack Training ran largely on Microsoft Azure using Microsoft Foundry with roughly 430 AMD MI300X GPUs in the cloud. For on-premises inference, AT&T and partners use AMD Instinct GPUs and the open ROCm software stack, avoiding CUDA dependency. Dell supplies carrier-grade servers equipped with AMD MI355X GPUs for local deployments. This combination provides a reproducible production blueprint for operators that need local control over sensitive workloads.
# Operational architecture: AI Gateway and routing
# Use cases in carrier operations OTel 2.0 is already applied to practical tasks: summarizing complex standards documents, producing compliant network configurations, assisting troubleshooting, creating runbooks, and retrieving technical knowledge. Because the training data explicitly includes standards and engineering material, the model is better suited to produce outputs that match telecom requirements.
# Security and validation needs Open models can introduce new attack vectors, such as prompt injection and model poisoning. Telecom networks require strict validation before automation touches live systems. AT&T publishes model weights and documentation on Hugging Face and plans weekly updates to support operator review and ongoing validation.
# Vendor and operational considerations The deployment demonstrates that large telecom AI workloads can run without CUDA/NVIDIA, but it does not remove vendor dependencies entirely. Operators must evaluate long-term support, integration, and maintenance for AMD-based stacks, cloud platforms used during training, and on-premises hardware. The production blueprint relies on engineering and volume to reach the cost efficiencies AT&T reports.
# What this means for operators OTel 2.0 gives a concrete example of running a domain-specialized model in production with on-premises inference and a hybrid routing approach. Operators evaluating Telco AI should compare dataset provenance, hardware/software stacks, integration effort, validation processes, and expected traffic volumes before aiming for similar cost and operational results.