AT&T Launches OTel 2.0 Open-Source AI for Telecom Networks

AT&T Launches OTel 2.0 Open-Source AI for Telecom Networks

Vladislav Zaimov stands at the forefront of a pivotal shift in how global networks are managed and secured. With a distinguished career focused on enterprise telecommunications and the complex risk management of vulnerable infrastructure, he has witnessed the transition from rigid hardware to the fluid, AI-driven architectures of 2026. His deep involvement in large-scale system deployments provides him with a unique vantage point on how open-source intelligence is currently redefining the operational reality for carriers worldwide.

OTel 2.0 utilizes the Google Gemma 4 31B-IT base and 400 billion telecom-specific tokens. Why was this 31-billion parameter size chosen over larger frontier models, and how does the specialized training data improve performance on 3GPP and O-RAN standards?

The decision to utilize a 31-billion parameter model like Gemma 4 is a strategic move toward operational sovereignty and efficiency. While frontier models are impressive, their massive size often requires an external cloud dependency that most carriers want to avoid for core network functions. By focusing on a refined pool of 400 billion tokens—narrowed down from an initial one trillion—we have a model that is surgically precise regarding 3GPP, ETSI, and O-RAN specifications. This smaller footprint allows the model to run comfortably on-premises within an operator’s own data center, ensuring that sensitive engineering data never leaves the building. We are seeing much higher accuracy in technical tasks because the model isn’t “distracted” by general consumer data, focusing instead on the dense protocols required for modern connectivity.

This project utilizes AMD Instinct GPUs and the open ROCm software stack rather than the standard NVIDIA/CUDA ecosystem. What technical challenges arise when shifting to this hardware for large-scale training, and how do Dell’s carrier-grade servers facilitate on-premises deployments for sensitive network tasks?

Shifting away from the dominant CUDA ecosystem requires a rigorous commitment to the ROCm open software stack, which involves re-validating training pipelines to ensure they perform at scale. During the development of OTel 2.0, we utilized approximately 430 AMD MI300X GPUs via Microsoft Azure for the heavy lifting of training, proving that massive workloads can thrive outside the traditional hardware monopolies. For the localized, day-to-day operations, Dell’s carrier-grade servers equipped with AMD MI355X GPUs provide the ruggedized, high-performance environment necessary for sensitive network tasks. These servers are specifically designed for the high-availability requirements of a telecommunications hub, offering a clear blueprint for operators who need to balance AI power with strict local control and data privacy.

The AI Gateway architecture uses cached routing to direct tasks between OTel 2.0 and larger models, reportedly cutting costs by up to 90%. How does this routing logic determine which tasks stay local, and what specific engineering resources must a smaller operator possess to replicate these types of savings?

The AI Gateway acts as an intelligent air traffic controller, using a caching layer to identify repetitive queries and routing them to the most efficient resource. If a request involves standard troubleshooting or summarizing a known ITU or CAMARA document, the logic keeps it local on OTel 2.0, which is where the 90% cost reduction primarily comes from. More abstract or highly complex creative tasks are routed to larger frontier models, but only when necessary. For a smaller operator to replicate this, they can’t just buy the hardware; they need a dedicated team of site reliability engineers and AI specialists who can fine-tune these routing thresholds based on their specific traffic volume. Without that high-level engineering depth, the overhead of managing such a complex gateway could quickly eat into the projected savings.

Open-source models in telecom environments introduce unique security risks like prompt injection and model poisoning. What specific validation steps are required before an AI can safely generate network configurations or automate troubleshooting, and how do weekly updates on Hugging Face help or hinder that security audit process?

Before any AI-generated configuration touches a live system, it must pass through a multi-stage validation sandbox that checks for syntax errors and protocol compliance against current standards. We treat the model’s output as “untrusted code” until it is verified by automated testing scripts that simulate the network environment. The weekly updates we provide on Hugging Face are a double-edged sword; they ensure that the community can patch vulnerabilities and update model weights rapidly, but they also mean the platform is a moving target. This requires a continuous security audit process where each new iteration is scanned for poisoning or backdoors before it is integrated into the production environment. It shifts our perspective from a “set it and forget it” mentality to one of constant, vigilant oversight.

OTel 2.0 is designed to summarize dense engineering materials and create operational runbooks. Could you walk through a step-by-step example of how a technician would use this model during a live network outage, and what metrics would you use to measure the model’s accuracy compared to human knowledge retrieval?

Imagine a technician facing a critical failure in a 5G core segment; instead of digging through thousands of pages of TM Forum documentation, they ask OTel 2.0 to summarize the specific error codes. The model instantly retrieves the relevant engineering standards and generates a step-by-step runbook tailored to that specific vendor’s hardware configuration. To measure success, we track the “Time to Knowledge Retrieval” and the “Actionable Accuracy” of the generated steps compared to a senior engineer’s manual process. We have found that the model significantly reduces the “fog of war” during the first fifteen minutes of an outage, which is the most critical window for preventing a total system collapse. It’s not just about speed; it’s about the sensory relief of having a clear, accurate path forward when the pressure is at its peak.

What is your forecast for the adoption of open-source, domain-specific AI models within the global telecommunications industry?

I expect that starting in 2026 and continuing through 2028, we will see a decisive move away from generic AI toward these “expert-in-a-box” models that live within the operator’s firewall. The industry is rapidly moving past the trial phase and into an operational reality where domain-specific models like OTel 2.0 become as standard as the routers and switches they manage. This shift will effectively break the vendor lock-in that has plagued telecom for decades, allowing even mid-sized operators to run sophisticated, automated networks with a fraction of the previous administrative overhead. Ultimately, the winners in this space will be those who embrace the transparency of open-source weights while maintaining the rigorous security standards required for critical infrastructure.

Subscribe to our weekly news digest.

Join now and become a part of our fast-growing community.

Invalid Email Address
Thanks for Subscribing!
We'll be sending you our best soon!
Something went wrong, please try again later