
Letting Data Speak, AI Act!
Case Study
Data ScienceAI-Powered Web-RTC Meeting System
Overview
An enterprise AI-powered video training platform provider needed a scalable, compliant infrastructure to support real-time WebRTC communications and multiple AI agents for sales and customer success training scenarios. We designed and implemented a multi-AZ Amazon EKS cluster using Terraform IaC, deploying LiveKit servers alongside TTS, STT agents with comprehensive auto-scaling capabilities, resulting in a production-ready POC with zero public exposure, complete AWS service integration through VPC endpoints, and dynamic scaling that optimizes costs while maintaining performance during variable training loads.

About the Client
An enterprise-grade AI-powered video training platform provider serving organizations in sales enablement, partner training, and customer success domains.
The Challenge
The client needed to evaluate and implement LiveKit, an open-source WebRTC-based real-time communication platform, on a production-ready infrastructure to support their AI-powered video training solution. The existing infrastructure lacked:
- Scalable WebRTC Infrastructure: The platform required a highly scalable, enterprise-grade WebRTC infrastructure capable of handling multiple concurrent real-time video sessions for training and assessment scenarios.
- Multi-Agent Architecture Support: The solution needed to seamlessly run multiple AI agents including Text-to-Speech (TTS), Speech-to-Text (STT), Large Language Models (LLMs), and the client's proprietary training assessment agents, all working in coordination.
- Regulatory Compliance Through Data Locality: Enterprise customers demanded strict data residency requirements, necessitating infrastructure that could maintain data within specific AWS regions to meet regulatory compliance standards.
- Auto-Scaling Capabilities: The training platform experienced variable load patterns based on organizational training schedules, requiring intelligent auto-scaling for both compute resources and application pods to optimize costs while maintaining performance.
- Private Network Architecture: Security requirements mandated a private nodegroup for redis with no public endpoints, while still maintaining operational efficiency through secure access patterns and AWS service integrations.
Without a properly architected solution, the platform would face scalability bottlenecks during peak training periods, potential regulatory compliance violations, excessive infrastructure costs, and inability to deliver reliable real-time AI-powered training experiences.
Key Results
- Successfully deployed a production-ready Proof of Concept (POC) for LiveKit platform on a multi-AZ Amazon EKS cluster with comprehensive auto-scaling capabilities across three availability zones
- Implemented Infrastructure as Code using Terraform, enabling repeatable deployments and reducing infrastructure provisioning time by 85%
- Configured Horizontal Pod Autoscaler (HPA) and Cluster Autoscaler to dynamically scale LiveKit servers and AI agent pods based on CPU utilization, memory consumption, and active session metrics
Our Solution

The solution involved designing and implementing a comprehensive, enterprise-grade private EKS cluster architecture on AWS to host the LiveKit WebRTC platform and multiple AI agents for the video training platform:
High-Level Architecture
Infrastructure Architecture and Design
- Designed a multi-availability zone EKS cluster architecture spanning three availability zones (us-east-1a, us-east-1b, us-east-1c) to ensure high availability and fault tolerance for real-time video communications.
- Configured VPC networking with separate private subnets for EKS worker nodes (10.0.0.0/20, 10.0.16.0/20, 10.0.32.0/20) and LiveKit/Agent workloads (10.0.48.0/20, 10.0.64.0/20, 10.0.80.0/20), implementing network segmentation for enhanced security.
- Implemented NAT Gateways in each availability zone to enable outbound internet connectivity for private subnets while maintaining inbound isolation.
EKS Cluster Deployment with Terraform IaC
- Provisioned the entire EKS cluster infrastructure using Terraform Infrastructure as Code scripts, enabling version-controlled, repeatable deployments and facilitating future multi-region expansions.
- Implemented comprehensive IAM roles and Kubernetes Role-Based Access Control (RBAC) policies to enforce least-privilege access principles across the infrastructure.
- Set up AWS Secrets Manager integration for secure management of LiveKit API keys and agent configuration secrets.
LiveKit Platform Implementation
- Deployed the livekit component using helm charts.
- Deployed LiveKit server components on dedicated node groups with node taints (workload=livekit) to ensure workload isolation and optimal resource allocation.
- Configured Internal Network Load Balancer (NLB) for LiveKit services, exposing ports 7880 (HTTP/WebSocket) and 7881 (TCP/WebRTC).
- Implemented session routing and load balancing strategies to distribute WebRTC connections across multiple LiveKit server pods for optimal performance.
- Optimized WebRTC configuration parameters for low-latency video streaming and efficient bandwidth utilization during training sessions.
Multi-Agent Deployment Architecture
- Deployed Livekit TTS(Text-to-speech) and STT(Speech-to-text) agents.
- Configured pod-to-pod communication between LiveKit servers and AI agents through Kubernetes services and network policies for secure, low-latency data exchange.
- Set initial replica counts of 2 pods per agent type across multiple availability zones for high availability.
Comprehensive Auto-Scaling Configuration
- Implemented Kubernetes Horizontal Pod Autoscaler (HPA) for BongoLearn agents with custom scaling triggers: CPU utilization >70%, Memory utilization >75%, and custom metrics based on active training sessions.
- Deployed Cluster Autoscaler to automatically provision additional EC2 instances when pod scheduling fails due to insufficient cluster capacity, with node group configurations: Node Group 1 (Min=2, Desired=3, Max=5 t3.xlarge instances) for general workloads and Node Group 2 (Min=1, Desired=2, Max=4 c5.2xlarge instances) for LiveKit workloads.
- Established scaling policies that balance cost optimization during low-usage periods with performance requirements during peak training schedules.
Security Implementation and Hardening
- Configured Kubernetes Network Policies to restrict pod-to-pod communication to only authorized services, implementing micro-segmentation within the cluster.
- Enabled TLS/SSL encryption for all inter-service communications using AWS Certificate Manager and Kubernetes ingress controllers.
- Implemented security groups with strict inbound/outbound rules: LiveKit/Agent security group restricting access to authorized CIDR ranges and necessary ports only.
Related Case Studies
← Back to All Case Studies
Data Science
AI-Assistance for Customer Support for Rent-to-Own Industry
A rent-to-own industry organization struggled with inconsistent customer support quality and slow response times that impacted lead conversion rates. By implementing an AI-powered chat assistance system using AWS Bedrock and retrieval-augmented generation, the organization enabled agents to receive three context-aware response suggestions within seconds during live conversations. The solution leverages historical successful conversations through semantic search and Claude Haiku 4.5, ensuring every agent delivers high-quality, proven communication strategies regardless of experience level. The serverless architecture processes thousands of requests monthly while maintaining reliability through intelligent fallback mechanisms and comprehensive monitoring.
Read More
Data Science
Artificial Intelligence - Driven Candidate Screening Revolution
JashDS revolutionized a company's hiring process by developing a GenAI-powered candidate screener that reduced time-to-hire by 50% and improved hiring outcomes. The solution leverages advanced language models to conduct dynamic, role-specific interviews, automatically generating and adapting questions based on job descriptions and candidate responses.
Read More
Data Science
Artificial Intelligence Model for Retail Shelf Monitoring
JashDS revolutionized retail shelf management for a major grocery chain by developing an AI-powered real-time monitoring system. The solution utilized advanced computer vision techniques and deep learning models to detect out-of-stock and misplaced products, significantly improving inventory accuracy and enhancing the customer shopping experience while reducing manual labor costs.
Read MoreHave a similar challenge?
Connect with us
