Artificial intelligence is changing application workloads, but it has not changed what users expect. AI applications must remain available, responsive and secure as demand grows.
Behind every chatbot, copilot, recommendation engine, retrieval-augmented generation application and inference service is an application delivery layer that connects users and systems to the right resources.If that layer slows down or becomes unavailable, model quality matters little to the user.
As organizations expand their AI initiatives, Application Delivery Controllers (ADCs) can provide a resilient foundation for AI applications and services.
AI applications introduce infrastructure demands that conventional enterprise applications may not need, especially at the same scale.
Inference traffic can be unpredictable, with sudden bursts driven by user activity, automated processes or multiple applications consuming the same models. Requests may also vary significantly in their processing requirements, response times and use of backend compute resources.
At the same time, AI environments are becoming increasingly distributed. Models, inference services, APIs, data sources and storage platforms may run across on-premises infrastructure, public clouds, private clouds, containers or edge locations.
Despite these differences, the basic expectations remain familiar:
This makes application delivery an essential part of the AI architecture, not simply an infrastructure consideration added after deployment.
An ADC provides an intelligent control point between users or applications and the services processing their requests. An AI environment includes inference servers, AI APIs, model-driven microservices, data pipelines and other application components.
By inspecting service availability and distributing requests across suitable resources, an ADC can help prevent individual servers or endpoints from becoming overloaded. It can also redirect traffic when a service becomes unavailable or begins to experience degradation.
The Progress® Kemp® LoadMaster® solution provides Layer 4 through Layer 7 application delivery, advanced health checks, application-aware traffic management and support for hybrid, multicloud, edge and Kubernetes-based environments.
These capabilities help organizations maintain a consistent delivery layer as AI services expand across different infrastructure models.
The ADC can support several important requirements for AI services.
AI applications can become integral to customer experience, employee productivity, analytics and automated business processes. When an inference endpoint or supporting service fails, the ADC can stop directing traffic to the affected resource and send requests to healthy alternatives.
Advanced health checks can identify both service failures and degradation, helping infrastructure teams maintain access to applications even when individual components experience problems. The LoadMaster solution also supports high-availability configurations and Global Server Load Balancing for resilience across multiple sites, data centers and cloud environments.
AI requests are not always uniform. Some may require greater processing capacity, access to a specific API or a connection to a particular service location.
Application-aware traffic management allows delivery decisions to consider factors such as endpoint health, service behavior and application requirements. The LoadMaster ADC can route and optimize traffic across AI workloads and inference endpoints while using health information to avoid unavailable or degraded services.
This helps organizations use their infrastructure more effectively while supporting a consistent experience during periods of variable demand.
AI services place unique demands on infrastructure because not all backend resources have the same capacity available at any given moment. A server may be healthy and online, but already operating near its performance limits due to high GPU, CPU or memory utilization.
In these environments, simply distributing traffic across available servers may not deliver the best user experience. Advanced application delivery solutions can make routing decisions based on real-time resource consumption, directing requests to backend systems with the most available capacity.
This resource-aware approach helps optimize AI inference performance, prevent resource bottlenecks, improve application responsiveness, and maximize the utilization of expensive GPU infrastructure. By considering workload conditions alongside service availability, organizations can deliver more consistent performance as AI demand scales.
4. Resilience Across Hybrid and Multicloud Environments
AI rarely operates in a single environment. An organization might train models in the cloud, run sensitive inference workloads on-premises, deploy services at the edge and make capabilities available through containerized APIs.
This distributed model creates an application delivery challenge. Users and systems still need reliable access regardless of where each component runs.
The LoadMaster solution supports deployment across hardware, virtualized infrastructure, public and private clouds, multicloud environments, Kubernetes and edge locations. This gives organizations the flexibility to establish more consistent application delivery policies without redesigning the delivery architecture for every environment.
AI services often rely heavily on APIs, making those interfaces an important part of the organization’s attack surface.
Placing security controls at the application delivery layer can help organizations inspect and manage requests before they reach backend services. The LoadMaster solution includes capabilities such as Web Application Firewall and protection, authentication, access controls, TLS management and traffic inspection for applications and APIs.
These controls are not a substitute for a complete AI security and governance strategy. However, they can strengthen the delivery path and help reduce the exposure of AI applications and supporting APIs.
As interest in enterprise AI grows, some vendors are combining application delivery, model access, AI governance, security, cost management and infrastructure management into broad platforms.
For certain organizations, this platform approach may be appropriate. For others, it can introduce capabilities, integrations and operational responsibilities that exceed their immediate requirements.
Before selecting an AI application delivery strategy, infrastructure teams should consider several practical questions:
The goal should not be to acquire the largest possible collection of features. It should be to build an architecture that meets operational requirements without creating unnecessary complexity.
A focused ADC strategy enables organizations to strengthen the availability, performance, resilience, and protection of AI services while continuing to use the security, governance, observability and model-management tools that best suit their environment.
AI models and supporting technologies will continue to evolve. Infrastructure strategies will change as organizations scale from initial proofs of concept to production applications.
The application delivery foundation must evolve with these technologies and deployment models.
That foundation should support services wherever they run, intelligently distribute traffic, respond to failures, protect exposed applications and APIs and provide the operational visibility needed to maintain reliable performance.
The Progress Kemp LoadMaster load balancer helps organizations deliver highly available and secure applications, APIs, and AI infrastructure across hybrid cloud environments. With flexible deployment options, advanced traffic management, health checking, security capabilities and centralized operational insights through LoadMaster 360, organizations can support AI initiatives without turning every infrastructure requirement into a broad platform purchase.
Build a More Resilient Delivery Foundation for AI
Explore how the Progress Kemp LoadMaster solution supports intelligent traffic management, high availability and application protection for AI workloads and APIs across hybrid and multicloud environments.
Start a Trial Today!
more from the author