Learn more about Riot Games, the company behind this role.
Open Roles
Principal Machine Learning Engineer - League of Legends
ML engineers at Riot own the full lifecycle of machine learning in production — from problem framing through model design, deployment, and operation — building systems that serve players at global scale. They work across disciplines with engineers, data scientists, designers, and product teams to turn ML capabilities into player-facing experiences and tools for making great games. The Role As a Principal AI/ML Platform Engineer for League of Legends , you will set ML platform and integration strategy that enables League teams to build, launch, and operate player and developer facing ML experiences. You will drive the systems, tooling, and operational standards that enable teams to deploy and operate ML at global scale. Your technical direction will influence product strategy and your work will directly shape how League of Legends builds and evolves player-facing experiences. Your work will span year-plus efforts across multiple products, mentoring senior engineers and coaching across disciplines. You will report to the Senior ML Engineering Manager within the League of Legends game team. This role will be based out our Los Angeles headquarters. Responsibilities - Lead AI/ML workflow and MLOps for League of Legends; set automation standards including auto-remediation and self-healing pipelines; drive implementation. - Lead AI/ML serving architecture that scales to production load for Riot’s player base; design systems requiring minimal operational intervention; drive implementation of proven serving architectures. - Lead ML feature platform in partnership with data engineering to scale feature development and processing; solve scale, latency, or reliability constraints in ML data pipelines. - Lead AI/ML developer experience that accelerates development; remove workflow bottlenecks and enable new capabilities. - Lead AI/ML governance in partnership with compliance experts to ensure regulatory adherence; drive implementation of proven frameworks. - Lead AI/ML platform security in partnership with security teams; drive implementation of proven privacy, security, and cryptography techniques for AI/ML. - Drive AI/ML pipeline development and deployment standards, coordinating with data engineering on shared MLOps capabilities. - Drive AI/ML service development and API standards across product integrations. - Drive operational excellence and cost optimization for AI/ML systems across the organization. - Drive incident management and response standards for AI/ML systems across production services. - Drive observability and monitoring standards for AI/ML systems across services. - Drive tooling strategy and evaluation for AI/ML platforms across the organization. - Drive mentorship of senior engineers across multiple products or problem areas; coach developers across disciplines. - Drive recruiting standards across multiple products or problem areas; contribute to interview kits and TA efforts. Qualifications - Bachelor’s degree in Computer Science or a related field, or equivalent practical experience. - 10+ years of professional software engineering experience, including 5+ years building production ML platforms, systems, or MLOps capabilities. - Evidence that platforms or systems you’ve led have become load-bearing for multiple teams or products — adopted, depended on, and evolved beyond the original scope. - Deep operational intuition for ML systems at scale: you’ve lived through the failure modes (training-serving skew, silent degradation, cost spirals) and built the systems to prevent them. - Background in cloud-native orchestration and large-scale system design (Kubernetes, GPU scheduling, container orchestration) applied at global scale. - History of improving developer experience for ML practitioners in ways that measurably changed how fast or reliably they shipped. - Track record mentoring senior and staff-level engineers; evidence of elevating platform engineering judgment across an organization. - Experience building ML platforms that serve multiple products or game titles simultaneously is a plus. - Track record solving novel scale, latency, or reliability constraints in ML data pipelines is a plus. - Background in ML governance, security, privacy-preserving techniques, or compliance at organizational scale is a plus. - Passion for player experience, games, or creative technology. For this role, you will find success through craft expertise, a collaborative spirit, and decision-making that prioritizes the delight of players. We will be looking at your past studies, experience, and your personal relationship with games. If you embody player empathy and care about players' experiences, this could be your role! Our Perks: Riot focuses on work/life balance, shown by our open paid time off policy and other perks such as flexible work schedules. We offer medical, dental, and life insurance, parental leave for you, your spouse/domestic partner, and children, and a 401k with company match. Check out our benefits pages for more information. At Riot Games, we put players first . That mission drives every decision in our quest to create games and experiences that make it better to be a player. Whether you’re working directly on a new player-facing experience or you’re supporting the company as a whole, everyone at Riot is part of our mission. And just like in our games, we’re better when we work together. Our goal is to create collaborative teams where you are empowered to bring your unique perspective everyday. If that sounds like the kind of place you want to work, we’re looking forward to your application. It’s our policy to provide equal employment opportunity for all applicants and members of Riot Games, Inc. Riot Games makes reasonable accommodations for handicapped and disabled Rioters and does not unlawfully discriminate on the basis of race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, handicap, veteran status, marital status, criminal history, or any other category protected by applicable federal and state law. We consider for employment all qualified applicants, including those with criminal histories, in a manner consistent with applicable federal, state and local law, including the California Fair Chance Act, the City of Los Angeles Fair Chance Initiative for Hiring Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, the San Francisco Fair Chance Ordinance, and the Washington Fair Chance Act. Per the Los Angeles County Fair Chance Ordinance, the following core duties may create a basis for disqualifying candidates with relevant criminal histories: - Safeguarding confidential and sensitive Company data - Communication with others, including Rioters and third parties such as vendors, and/or players, including minors - Accessing Company assets, secure digital systems, and networks - Ensuring a safe interactive environment for players and other Rioters These duties are directly related to essential operations, safety, trust, and compliance obligations within our organization. Please note that job duties may evolve based on business needs and additional responsibilities may be assigned as necessary to maintain operational efficiency and security. (Los Angeles Only) Base salary range between $292,300.00 - $437,900.00 USD + incentive compensation + equity + 401K with company match + medical, dental, vision, and life insurance + short and long-term disability + open PTO.
Senior Software Engineer, Platform & Infrastructure - Riot Technology
Platform and infrastructure engineers at Riot build the foundational systems that enable teams to develop, deploy, and operate systems at global scale. They partner across disciplines with other engineers, data scientists, designers, and product teams to ensure reliable, scalable, and secure infrastructure underpins every ML capability that reaches players. As a Senior Platform & Infrastructure Engineer on the Riot Technology team, you will design, build, and operate the core infrastructure and ML platforms. Your focus will be on the computing and orchestration platforms that power large-scale distributed training of agents (e.g. via RL, IL, and other techniques), simulation environments, and policy evaluation, as well as the CI/CD, infrastructure-as-code, observability, and developer tooling that keep these systems production-grade. You will close critical infrastructure gaps across the team's stack, driving improvements to standards, automation, and operational maturity. You will operate independently on multi-month work efforts and begin to influence technical direction beyond your immediate team. You will report to the Manager of Machine Learning. Responsibilities: - Build and operate Kubernetes, multi-node GPU clusters, and networking infrastructure for distributed ML bot training and large-scale policy evaluation. - Design infrastructure for running simulation environments at scale, enabling parallel rollouts, data collection, training, and evaluation. - Build CI/CD, deployment automation, artifact management, and infrastructure-as-code across cloud environments. - Improve platform reliability, cost efficiency, performance, reproducibility, auditability, and operational maturity. - Build observability, monitoring, alerting, health indicators, and SLO-aligned dashboards for infrastructure and ML workloads. - Develop internal APIs, control planes, templates, and developer tooling for distributed training and evaluation workflows. - Support MLOps workflows including automated training pipelines, model artifact management, experiment tracking, and reproducible ML lifecycle operations. - Build security and governance controls, manage production incidents, drive root-cause remediation, mentor engineers, and support recruiting for platform roles. Required Qualifications: - Bachelor’s degree in Computer Science or a related field, or equivalent practical experience. - 3+ years of software engineering experience, with meaningful experience in infrastructure, platform engineering, or SRE roles. - Experience operating distributed systems in production and keeping them healthy under real load. - Strong experience with Kubernetes, AWS or GCP, infrastructure-as-code, CI/CD, deployment automation, and production tooling. - Experience with GPU compute infrastructure, including scheduling, multi-node orchestration, and resource optimization for long-running training workloads. - Proficiency in Python and solid understanding of networking, microservices, core infrastructure services, and distributed systems fundamentals. Desired Qualifications: - Familiarity with MLOps workflows such as model versioning, pipeline orchestration, experiment tracking, artifact management, and reproducible ML workflows. - Experience with distributed training or HPC frameworks, inference serving, systems languages, high-performance networking, Unreal/client-server architecture, AI-assisted development tools - Passion for games and player experience. For this role, you'll find success through craft expertise, a collaborative spirit, and decision-making that prioritizes the delight of players. We will be looking at your past studies, experience, and your personal relationship with games. If you embody player empathy and care about players' experiences, this could be your role! Our Perks: Riot focuses on work/life balance, shown by our open paid time off policy and other perks such as flexible work schedules. We offer medical, dental, and life insurance, parental leave for you, your spouse/domestic partner, and children, and a 401k with company match. Check out our benefits pages for more information. At Riot Games, we put players first . That mission drives every decision in our quest to create games and experiences that make it better to be a player. Whether you’re working directly on a new player-facing experience or you’re supporting the company as a whole, everyone at Riot is part of our mission. And just like in our games, we’re better when we work together. Our goal is to create collaborative teams where you are empowered to bring your unique perspective everyday. If that sounds like the kind of place you want to work, we’re looking forward to your application. It’s our policy to provide equal employment opportunity for all applicants and members of Riot Games, Inc. Riot Games makes reasonable accommodations for handicapped and disabled Rioters and does not unlawfully discriminate on the basis of race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, handicap, veteran status, marital status, criminal history, or any other category protected by applicable federal and state law. We consider for employment all qualified applicants, including those with criminal histories, in a manner consistent with applicable federal, state and local law, including the California Fair Chance Act, the City of Los Angeles Fair Chance Initiative for Hiring Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, the San Francisco Fair Chance Ordinance, and the Washington Fair Chance Act. Per the Los Angeles County Fair Chance Ordinance, the following core duties may create a basis for disqualifying candidates with relevant criminal histories: - Safeguarding confidential and sensitive Company data - Communication with others, including Rioters and third parties such as vendors, and/or players, including minors - Accessing Company assets, secure digital systems, and networks - Ensuring a safe interactive environment for players and other Rioters These duties are directly related to essential operations, safety, trust, and compliance obligations within our organization. Please note that job duties may evolve based on business needs and additional responsibilities may be assigned as necessary to maintain operational efficiency and security.
Principal Software Engineer - DevOps / Site Reliability Engineer
Riot Games was established in 2006 by entrepreneurial gamers who believe that player-focused game development can result in great games. In 2009, Riot released its debut title League of Legends to critical and player acclaim. As the most played PC game in the world, over 100 million play every month. Players form the foundation of our community and it’s for them that we continue to evolve and improve the League of Legends experience. We’re looking for humble but ambitious, razor-sharp professionals who can teach us a thing or two. We promise to return the favor. Like us, you take play seriously; you’re passionate about games. We embrace those who see things differently, aren’t afraid to experiment, and who have a healthy disregard for constraints. That's where you come in. The AI Efficiency team at Riot Games builds the platforms, tools, and technical foundations that help Rioters safely and effectively use AI to accelerate how we work. As these systems become increasingly important to creative, product, and development workflows across Riot, we need dedicated engineering leadership to ensure these systems remain stable, scalable, secure, and dependable in production. As a Principal DevOps / Site Reliability Enginee r on the AI Efficiency team, you will own and evolve the operational foundations that allow the AI Efficiency team’s tech platform and the tools deployed within it to run reliably at growing scale. You will establish the systems, standards, automation, and support practices required to move quickly without compromising availability, deployment safety, maintainability, or user trust. You will partner closely with software engineers, ML platform engineers, technical artists, data scientists, and Riot’s infrastructure and security teams to improve developer experience, production readiness, observability, incident response, capacity planning, and service resilience. You will also help evaluate and operationalize AI-native engineering workflows such as agent-assisted code review, automated bug triage, AI-driven performance and security analysis, and browser-based UI validation. This role ensures the broader platform and its services are safely operated, supported, and continuously improved in production. You’re right for this role if you enjoy making complex systems reliable, reducing operational toil, improving how engineers build and ship software, and anticipating how systems will fail before those failures affect users. You are comfortable taking ownership of production health, leading through incidents, building sustainable operational practices, and creating paved roads that help teams move quickly and safely. You are also energized by the opportunity to responsibly bring new AI-native automation patterns into real engineering workflows, thoughtfully applying emerging capabilities to reduce friction, improve reliability, and enhance how engineers interact with production systems without compromising safety or control. Responsibilities: - Own and continuously improve the reliability, availability, scalability, performance, and operational health of the Efficiency team’s (web) platform and the tools deployed within it - Design, build, and maintain the infrastructure, deployment systems, and operational foundations required to support a growing portfolio of production AI services and internal tools - Improve CI/CD pipelines, release engineering practices, environment management, and deployment automation so software can be shipped safely, quickly, and consistently - Establish production-readiness standards and ensure new utilities have appropriate monitoring, alerting, ownership, documentation, rollback strategies, and support plans before launch - Define and operationalize service health indicators, SLIs, SLOs, error budgets, and reliability metrics that guide engineering priorities and tradeoffs between reliability, velocity, cost, and complexity - Build comprehensive observability across applications, infrastructure, service dependencies, and user workflows using metrics, logs, traces, dashboards, synthetic monitoring, and actionable alerts - Establish sustainable incident-management and on-call practices, including escalation paths, runbooks, severity definitions, communication protocols, and clear service ownership - Lead or contribute to the diagnosis and resolution of production incidents, coordinating across teams and driving blameless post-incident reviews and durable corrective actions - Build automation that reduces operational toil, improves mean time to detect and recover, and eliminates recurring sources of failure or manual intervention - Implement safe deployment patterns such as automated validation, progressive delivery, canary releases, feature flags, health checks, rollback mechanisms, and controlled environment promotion - Perform capacity planning, load testing, performance analysis, and resource forecasting to ensure the Toolkit can support increasing adoption and usage across Riot - Design and validate resilience, backup, recovery, failover, and disaster-recovery strategies for critical services, data, configurations, and infrastructure - Identify single points of failure and systemic risks across applications, cloud infrastructure, networking, databases, queues, caches, third-party dependencies, and operational workflows - Improve developer experience by building self-service workflows, reusable infrastructure components, local development environments, test environments, deployment tooling, and clear operational documentation - Establish and maintain infrastructure-as-code, configuration-management, secrets-management, and environment-governance practices that make infrastructure changes safe, repeatable, and auditable - Partner with engineers throughout the software development lifecycle to embed reliability, operability, security, and maintainability into system design rather than addressing them only after launch - Troubleshoot complex production issues across web applications, APIs, distributed services, containerized workloads, cloud infrastructure, network boundaries, authentication systems, and external service dependencies - Partner with ML Platform Engineers to ensure model-serving and inference systems integrate cleanly with the team’s broader observability, deployment, incident-management, and reliability standards - Collaborate with Riot infrastructure, information security, IT, developer-platform, and compliance teams to ensure the team follows appropriate operational and security requirements - Evaluate and implement AI-assisted operational workflows such as automated anomaly investigation, log analysis, remediation recommendations, regression detection, and runbook automation - Define guardrails, approval requirements, auditability, and escalation paths for agentic or automated operational systems that can interact with production environments - Champion operational excellence through technical leadership, mentoring, documentation, standards, architecture reviews, and tooling that raise the reliability bar across the team Required Qualifications: - Bachelor’s degree in Computer Science or a related field, or equivalent professional experience - 5+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, Platform Engineering, Production Engineering, Developer Experience, or a similar role supporting production systems - Strong programming and automation skills in one or more languages such as Python, Go, JavaScript, or TypeScript - Experience designing, operating, and improving cloud-based production systems in AWS, GCP, Azure, or comparable environments - Experience building and maintaining CI/CD pipelines, release systems, deployment automation, and environment-management workflows - Strong understanding of observability practices, including metrics, logging, distributed tracing, dashboards, synthetic monitoring, and alert design - Experience participating in or leading incident response, on-call support, root-cause analysis, and post-incident improvement work - Experience improving the reliability, availability, scalability, and performance of distributed systems, service-oriented architectures, APIs, or web platforms - Strong understanding of containerized environments and orchestration technologies such as ECS, Docker, Kubernetes, or comparable systems - Experience with infrastructure-as-code and configuration-management tools such as Terraform, Pulumi, CloudFormation, or similar technologies - Working knowledge of Linux systems, networking, DNS, load balancing, service discovery, authentication, secrets management, and cloud security fundamentals - Ability to identify systemic operational risks and drive durable improvements across systems owned by multiple engineers or teams - Ability to collaborate across organizational boundaries, influence technical direction, and communicate clearly during both planned work and high-pressure incidents - Experience providing technical leadership, mentoring engineers, and establishing engineering standards across a team or organization Desired Qualifications: - Experience supporting AI/ML platforms, inference services, model-serving systems, GPU-backed workloads, data pipelines, or other compute-intensive services - Experience defining and using SLOs, error budgets, and reliability metrics to guide prioritization and engineering decisions - Experience designing or improving internal developer platforms, self-service infrastructure, paved roads, golden paths, or shared engineering services - Experience building sustainable on-call rotations and operational support models for services used by multiple teams - Experience with progressive delivery, canary deployments, blue-green deployments, feature-flag systems, and automated rollback strategies - Experience with performance testing, capacity modeling, chaos engineering, fault injection, resilience testing, or failure-mode analysis - Experience designing backup, disaster-recovery, business-continuity, and regional failover strategies - Experience operating databases, caches, message queues, object storage, service meshes, API gateways, and other common distributed-system components - Experience improving security posture through access controls, secrets management, dependency management, vulnerability remediation, network segmentation, and infrastructure hardening - Experience balancing availability, latency, engineering velocity, infrastructure efficiency, and cost in systems operating at scale - Familiarity with browser automation and end-to-end testing frameworks such as Playwright for validating critical user workflows and detecting production regressions - Experience evaluating or integrating AI-assisted tools for incident investigation, anomaly detection, operational diagnostics, code review, test generation, or automated remediation - Familiarity with the risks and operational controls required when AI agents interact with source control, CI/CD pipelines, cloud infrastructure, or production systems - Experience establishing governance, approval workflows, audit trails, and quality controls for automated operational systems - Experience working in environments where experimental tools must be transitioned into reliable, supported, and maintainable production services For this role, you'll find success through craft expertise, a collaborative spirit, and decision-making that prioritizes your fellow Rioters, who are the customers of your work. Being a dedicated fan of games is not necessary for this position! Our Perks: - Full relocation support - Comprehensive health insurance for you, your spouse, and children - Open paid time off - Retirement benefits with company matching - Life insurance, parental leave, plus short-term and long-term disability - Play Fund so you can deepen your knowledge of our players and community through games - We’ll double down on your donations of time and money to non-profits
Company Details
Registered Agents
No registered agents are associated with this company yet.