AI adoption has reached a critical turning point, moving from training to the inference phase of deployment.. The challenge, however, no longer rests in proving the technology’s potential, but deploying it safely, reliably and at scale. As a result, many companies get bogged down in the pilot phase of AI, due to concerns around data protection, operational risk and the infrastructure required to support heavy workloads..
Experts from SDxCentral, Nokia, Gcore and DCD Intelligence come together in a recent DCD>Broadcast episode to explore how enterprises can overcome troublesome barriers as sovereignty and compliance requirements continue to evolve.. The ‘tokenomics’ of AI. Over the last several years, the industry has focused on building increasingly larger AI training models.
When building, success was often measured by how quickly the model was trained and how much training cost.. The pathway has now shifted.. “We’ve entered a new phase where value is created every time a model is used,” says Patrick McCabe, strategist for AI networking at Nokia.
“This leads us to the notion of inference and the token economy.”. Tokens are defined as basic units of information that an AI model processes. Every user request is converted into input tokens, with every response generated as output tokens.
When it comes to modern AI systems, much more data is processed and can represent a greater depth of material – information such as conversation history, agent instructions, and intelligence in general.. “Tokens have become the common currency of the AI economy” Patrick McCabe, Nokia. “A single user interaction can generate 1000s or even millions of tokens flowing through the distributed AI infrastructure,” McCabe explains.
“That’s why tokens have become the common currency of the AI economy. They represent the actual work performed by AI systems and represent the currency charged for the service.”. During a time when agentic AI systems are coming online, the industry expects token consumption to explode.
For McCabe, the growth is already staggering.. “It’s exciting and frightening at the same time,” he explains. “Some recent data showed growth measured in tokens per month will grow from five quadrillion today to over 120 quadrillion by 2030.
This is a massive new distributed compute economy and the network must support it.”. Building the foundation behind agentic workloads. With inferencing growing exponentially, many projects are centered around gigawatt data centers.
AI prompts are split into different components, often processed at different times and locations. This is what makes the network so critical, particularly as AI moves from isolated clusters to federated systems.. However, Seva Vayner, product director of Edge, cloud and AI at Gcore, explains that the rate of demand and consumption are outpacing market capabilities..
“With regulations coming in, you need to potentially distribute the load you have between the multiple sides where you can find the capacity,” he says.. McCabe explains how components are not located in the same place, but instead distributed across clouds, regional hubs and Edge infrastructure..
“The network has to coordinate these resources, move context between states and enable them to operate as a single system,” he says.. The end goal, he adds, is to have a single overarching system that understands all components within it and then optimizes multiple specific modular requests..
“In many ways, the future challenge of AI is no longer simply creating intelligence,” he notes. “It’s efficiently connecting and coordinating intelligence across the infrastructure.”. As a result of AI training within a single cluster, three main networking vectors have emerged: scale up, scale out and scale across.
This defines how the industry has been moving rapidly toward distributed compute and inferencing, moving beyond GPU clusters into distributed AI systems.. “These principles need to be applied to the new vector of networking,” McCabe explains. “It’s not just one technology; there’s a combination of networking technologies that need to be used – all play a huge role in providing the networking performance needed to satisfy the end user.”.
For Gcore, which is currently running multiple points of presence (PoPs) across different locations, starting with multiple layers to optimize the network is important. Vayner shares how the company is moving into the inference engine and logic optimization phase for AI models.. “Optimization starts at multiple layers, beginning with the infrastructure side,” he says.
“It’s also about how fast you can run and scale your models, especially for GPU weights. You don’t want to download those weights every time – you want to scale up and upload the weights into GPU memory, which takes time.”. Gcore’s background in building network infrastructure and shifting to a distributed AI platform helps it to scale out inference, moving from simply running models and processing user requests to enabling real-time interaction with models..
From Nokia’s perspective, this establishes greater trust when building out the network, particularly where AI workloads are concerned.. “Demand and consumption is growing so significantly, it’s more than the market can provide” Seva Vayner, Gcore. “It’s a complex environment behind the scenes.
Every millisecond matters,” McCabe says. “If the network introduces latency, congestion, packet loss or unpredictability, the AI experience may suffer.”. Consequences could involve longer response times or slower token generation, which could reach the end user – having potentially detrimental impacts in mission-critical environments..
“A world-class AI network needs to eliminate frictions and provide a deterministic level of performance,” he adds. “You need to be able to predict the performance your network can deliver across the entire system. When you know that, you can make decisions based on it..
“Rather than being a bottleneck, we like to think we’re successful when there’s an invisible fabric that allows distributed compute resources to behave as though they are part of a single, cohesive AI system.”. Building trust and sovereignty in a shared AI infrastructure. The AI landscape continues to shift away from so-called ‘plumbing’ toward inference, turbocharging the technology for optimization.
When it comes to the network, the wider market is having to evolve to meet expectations.. “Hyperscalers and API providers are already pricing in the pay-per-token, pay-per-call way, but I’m not sure if enterprise budgeting has caught up with that yet,” shares Barney Dixon, senior manager for insights and reports at DCD Intelligence..
Plenty of organizations continue to budget for AI as they budget for general infrastructure, but token consumption doesn’t work in the same way, Dixon explains.. “This makes enterprises nervous to sign up for these models,” he says. “Spend is unpredictable, especially with AI agents – enterprises will want a better understanding of consumption before they can commit to pricing models.”.
AI sovereignty and buying GPUs can make buyers nervous about their budget. With this in mind, the conversation considers multi-tenancy in connection with sovereignty as a trust consideration for AI distributed computing.. According to McCabe, Nokia has used multi-tenancy strategies with telcos for decades.
With each tenant having multiple services, its secure isolation operates at both the tenant and service layers.. “The reality is that tenants, services and applications must coexist on the same physical network,” McCabe explains. “In many cases, you have a shared infrastructure, which frightens people, but your environment remains separated from tenant to tenant in terms of traffic, performance, security and operations.”.
The optical layer can extend isolation further, providing a dedicated coherent wavelength for a tenant. It can leverage optical transport network (OTN) channels and paths, providing reserved bandwidth constructs that physically separate flows.. “The real value emerges when optical and routing layers work together,” McCabe says.
“It enables predictable performance, stronger security, the ability to provide regulatory compliance and the ability to assure SLAs that are in the tenant’s contract.”. This is now often the way for network transformation.. When looking at trust more broadly, beyond the technical side of scaling, many organizations are now working hard on isolation, encryption and auditability.
But while common industry practice, it’s not as simple as it seems.. “It’s a core focus,” Dixon explains. “The problem from the compliance side is that there’s no real standardization.”.
Securing the future. As the AI game advances, so does the digital threat landscape as the AI conversation shifts to include quantum-safe networking.. Concerns across the industry continue to mount over Q Day in particular, or Y2Q, the moment when quantum computers become powerful enough to break standard encryption.
For McCabe, this focus is the wrong approach.. “Plenty of organisations continue to budget for AI like they budget for general infrastructure, but token consumption doesn’t work in the same way” Barney Dixon, DCD Intelligence. “We shouldn’t be focusing on a future date; the risk window is already open,” he says.
“The idea is ‘harvest now, decrypt later’ – it’s expected that threat actors will be able to decrypt encrypted data they have stored in a few years’ time and then unleash havoc on the owner of the information.”. This risk is open to protected intellectual property like government information, financial records, healthcare data and customer information that must remain secure.
To protect it, McCabe says planning must take place now.. Attack surfaces are only growing, with distributed denial-of-service (DDoS) attacks growing more frequent and larger in aggregate. The recent hack on Hugging Face, for example, shows just how vulnerable digital systems can be to more sophisticated AI models..
“This is one of the biggest challenges right now, especially for security,” Vayner explains. “For overall DDoS attacks, we see botnets are growing significantly and, with the number of AI models and IoT devices increasing, the potential of huge infrastructure attacks is significant.”.
Targets of these types of attacks can impact AI training or inference too, which correlates with a company’s revenue plans or new funding rounds. For instance, increasing attack surfaces can make companies nervous to invest or proceed with their AI plans.. Gcore is already working on providing services to protect specific patterns and avoid incidents that are already happening across the market..
“It’s a challenging topic because models themselves will likely be protected by guardrail models that try to block harmful actions,” Vayner explains. “It will be an interesting new market with new innovations and technologies focused on protection.”. Building a stronger ecosystem.
As enterprises move from just consuming AI to operationalizing it within their own data, sovereignty has inevitably become a dominant part of the conversation.. “You might have your data stored in-house, but the model or infrastructure might call out to another country, which is an issue,” Dixon says.
“It’s about budgeting, but it’s also about data residency and access control, which is mostly seen in more regulated industries like finance and healthcare.”. With the power of AI today, it is evident that networking cannot be sold in isolation, but instead demands a strong ecosystem of partners and customers to build solutions together..
“You don’t just partner, you integrate with all of these entities across the entire stack, otherwise it’s not going to work,” McCabe explains.. For Nokia, this involves working with the likes of Gcore to test and integrate with server vendors, GPU vendors, NIC vendors and more, seeing this approach as a necessity to build systems as a collective..
“It’s no longer just about the network,” McCabe suggests. “We have to make sure we have strong partners and we integrate comprehensively and rigorously with them.”. For the full conversation, watch the full broadcast on-demand..
For more information, please visit: https://onestore.nokia.com/asset/215454?did=D00000014236&utm_campaign=Move_Fast_With_Confidence&utm_source=DCD&utm_medium=webinar&utm_content=article&utm_term=gcore
