Room S
Workshop
In person
LeadersGoldDiscovery

Agentic AI: Architecture and standards for next-generation AI Agents

  • Date
    9 July 2026
    Timeframe
    14:00 - 17:15 CEST
    Duration
    3h 15 minutes

      As Artificial Intelligence transitions from static models to autonomous agents capable of reasoning and planning, the industry requires robust technical foundations to support this shift. This workshop explores the emerging ecosystem of Agentic AI, with a focus on architecture and orchestration frameworks for multi-agent systems. We will explore interoperability protocols and standards to ensure safe, scalable, and transparent autonomous deployment across diverse technical environments.

      Beyond software frameworks, the session examines the convergence of AI and telecommunications, particularly how Agentic AI will shape next-generation AI-native networks (IMT-2030/6G). By analyzing the interplay between autonomous agents and network optimization, we will highlight use cases such as dynamic spectrum management for ultra-reliable low-latency communications, self-optimizing resource allocation in dense edge deployments, predictive traffic orchestration for massive IoT connectivity, and autonomous fault recovery in future network slicing architectures.

      Schedule

      This session gives an overview of key lessons learned from building agents in the real world. We start with background on AI and innovation at Google, from AlphaFold's impact on scientific discovery to multimodal foundation models like Gemini or Veo and emerging research agents that propose and test new ideas. From there we dive deep into real-world considerations. We focus on agents, which unlock new capabilities but also raise harder questions around reliability, safety, and trust. We will cover agent evaluation: how to measure quality, trajectory, and outcomes when behavior is non-deterministic, as well as responsible AI in practice.

      Many real-world diagnosis tasks can be viewed as deduction problems: a hidden cause must be inferred from incomplete observations. This challenge arises in both games and telecom network troubleshooting. In this talk I describe Simulation-Based AI (SBAI) methods that reason by using models of a problem domain rather than relying solely on learned patterns. Drawing on work in game AI, I show how these techniques can be combined with large language models to create more capable agentic systems, where the LLM orchestrates data extraction, network tool use, model-building, and problem-specific solvers to improve accuracy and speed of fault detection.

       

      Presentation slides

       

      Specialized agents can leverage the flexibility of LLMs to execute narrow tasks in dynamic feedback loops. This offers a powerful entry point to automate operational workflows that would otherwise be too labour intensive to conduct manually. However, such agents also need to operate reliably and predictably, while scaling to millions of requests at low cost. In practice, frameworks for assessing specialized agent reliability remain a major gap. This talk presents an introduction to testing and evaluating the reliability of specialized agents. With a practical grounding in governance workflows, we share lessons from testing agent reliability for consent verification and offer actionable insights to address this gap.

      AI is rapidly evolving from standalone models to autonomous agents capable of collaborating across applications, organizations, and networks. However, scalable multi-agent systems require standardized protocols for communication, coordination, trust, and interoperability. This talk explores the emerging protocol landscape for the agentic era, highlighting the role of intent-driven coordination as a bridge between high-level user goals and autonomous agent execution. Finally, it discusses why benchmarking agentic systems is becoming as important as benchmarking foundation models, and the need for new evaluation frameworks that measure coordination, reliability, trust, and real-world task performance.

      As AI agents evolve from conversational assistants into autonomous problem-solvers, they face a critical roadblock: how do they find the right agents, tools, skills, or APIs across a decentralized web at runtime? In this session, we introduce the Agentic Resource Discovery (ARD) specification, a new open standard that bridges the gap between AI agents and capabilities by standardizing how tools, skills, and agents are cataloged, searched, and discovered. During the session, we will explore the Agentic Resource Discovery specification hands-on to highlight its capabilities and discuss integration into existing landscapes.

      As AI agents move into telecom networks, how do you know which model to pick for a given task? How do you know when a model is safe enough to be deployed autonomously? Other domains have agentic benchmarks: SWE-Bench for coding, Cybench for cybersecurity, and HealthBench for healthcare. No comparable benchmark exists for telecoms. In this session, Enrique will provide a practical overview of the current state of AI evaluation, across research and practice; how the field has shifted from LLM benchmarks to agentic ones; and why evaluation must be a core part of AI development, not a final step in a training pipeline. Evaluating an agent is methodologically distinct from evaluating a model output. Enrique will preview NOC-Bench, a benchmark GSMA is building to evaluate AI agents as network engineers. The session will explore how agentic evaluation must become standard practice across the industry, and how operators, researchers, and vendors must agree on shared evaluation frameworks rigorous enough for regulated deployment. Attendees will leave with a clearer view of what agentic evaluation in telecoms has to measure, and where today's frontier models actually stand.

      AI agents and network sandboxes are poised to become fundamental building blocks of future AI-native telecommunication networks. This talk explores the strong and mutually beneficial relationship between these two technologies and how their synergy can drive networks toward greater efficiency, adaptability, and performance. Network sandboxes provide agents with realistic and controlled environments in which to learn, test, and refine their decision-making strategies before deployment in operational networks. In turn, agents can enhance the fidelity, scalability, and efficiency of sandbox. Beyond training and validation, sandboxes offer a valuable framework for benchmarking and comparing models and agents under diverse network conditions. This capability is particularly important for detecting performance degradation caused by evolving network environments, such as the introduction of new technologies, changing traffic patterns, or updated network configurations allowing service providers the opportunity to proactively identify when model retraining, adaptation, or replacement is required before significant performance losses occur. The talk will discuss the opportunities and challenges associated with building and maintaining this symbiotic relationship and will also present recent results demonstrating the impact of combining AI agents and network sandboxes in the development of robust, adaptive, and intelligent AI-native telecommunication networks.

      Zindi's first agentic AI challenge asked participants to go beyond simple prediction models and build agents that can actually find and fix network problems. Supported by Huawei, ITU, GSMA, and GenAINet, the challenge had teams build agents using simulation, memory, RAG, and skills, all on the same Qwen3.5 base model, across two tracks: wireless and IP network troubleshooting. This talk looks at how the challenge was designed: why everyone used the same base model, how the three-phase format pushed teams to build agents that generalise well (not just top the leaderboard), and what we learned about using agentic AI in real telecom settings.

       

      Presentation slides

       

      Diagnosing performance degradations in telecommunication networks is a time-consuming task that requires engineers to interpret large volumes of measurements, identify root causes, and determine appropriate corrective actions. Automating parts of this process has the potential to improve operational efficiency while reducing service impact. The Telco Troubleshooting Agentic Challenge, organized by the ITU, explores how AI agents can assist network engineers by analyzing network conditions and recommending appropriate troubleshooting actions. This presentation describes the solution I developed for the challenge, which combines feature engineering, machine learning, and LLMs within an agentic workflow. By integrating domain knowledge with data-driven decision-making, the solution shows how AI agents can support faster and more consistent network diagnosis while remaining interpretable and practical for real-world telecom operations.

       

      Presentation slides

       

      Share this session with your network

      Are you sure you want to remove this speaker?