Lead the deployment, operationalization, monitoring, and continuous improvement of AI/ML models in production. Establish MLOps practices, automation, reliability monitoring, and incident response processes. Partner with AI engineers, product teams, business stakeholders, risk, compliance, and governance teams to define success metrics, meet regulatory and ethical standards, resolve operational issues, and demonstrate business value.
Our Purpose
Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we're helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential.
Title and Summary
Lead Site Reliability Engineer (AI/ML)
As a Lead Site Reliability Engineer at Mastercard, you'll play a pivotal role focusing on the seamless deployment, operationalization, and continuous improvement of our AI/ML solutions. You'll be instrumental in translating AI models from development to production, ensuring they deliver tangible business value, operate efficiently, and meet key performance indicators.
Key Responsibilities• Lead the E2E deployment and operationalization of AI/ML models and solutions, ensuring they are scalable, reliable, and integrated seamlessly into existing business processes• Establish and maintain robust monitoring frameworks for deployed AI solutions. Proactively identify performance bottlenecks, data drifts, and other issues, and drive their resolution to ensure optimal business outcomes• Work closely with business stakeholders, AI Engineers, and product teams to understand business requirements, define success metrics for AI solutions, and ensure deployed models are directly contributing to key business objectives• Implement and champion MLOps best practices, automation strategies, and efficient workflows to streamline the deployment lifecycle of AI models, from experimentation to production• Collaborate with risk, compliance, and governance teams to ensure all AI deployments adhere to internal policies, regulatory requirements, and ethical AI principles• Lead the response to operational incidents related to deployed AI models, conducting root cause analysis and implementing preventative measures
Qualifications• Education: Bachelor's degree in Computer Science, Engineering, Data Science, Business, or a related field• Experience: Minimum of 8+ years of experience in AI/ML operations, MLOps, DevOps, or a related role with a strong focus on deploying and managing AI/ML solutions in production environments.• Technical Skills:
o Solid understanding of the AI/ML lifecycle, from data preparation and model training to deployment and monitoring.
o Experience with one of the cloud platforms and their AI/ML services
o Proficiency in scripting and
o Familiarity with containerization technologies
o Knowledge of CI/CD pipelines for machine learning models.
o Experience with monitoring tools for AI/ML solutions
o Understanding of data governance, data quality, and data security principles relevant to AI/ML• Strong ability to understand business needs, translate them into technical requirements for AI solutions, and articulate the business value of AI deployments• Excellent communication, interpersonal, and stakeholder management skills• Ability to effectively bridge the gap between technical and business teams• Demonstrated ability to lead initiatives, drive cross-functional projects, and influence outcomes without direct authority• Strong understanding of operational processes and a passion for optimizing them
Corporate Security Responsibility
All activities involving access to Mastercard assets, information, and networks comes with an inherent risk to the organization and, therefore, it is expected that every person working for, or on behalf of, Mastercard is responsible for information security and must:
Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we're helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential.
Title and Summary
Lead Site Reliability Engineer (AI/ML)
As a Lead Site Reliability Engineer at Mastercard, you'll play a pivotal role focusing on the seamless deployment, operationalization, and continuous improvement of our AI/ML solutions. You'll be instrumental in translating AI models from development to production, ensuring they deliver tangible business value, operate efficiently, and meet key performance indicators.
Key Responsibilities• Lead the E2E deployment and operationalization of AI/ML models and solutions, ensuring they are scalable, reliable, and integrated seamlessly into existing business processes• Establish and maintain robust monitoring frameworks for deployed AI solutions. Proactively identify performance bottlenecks, data drifts, and other issues, and drive their resolution to ensure optimal business outcomes• Work closely with business stakeholders, AI Engineers, and product teams to understand business requirements, define success metrics for AI solutions, and ensure deployed models are directly contributing to key business objectives• Implement and champion MLOps best practices, automation strategies, and efficient workflows to streamline the deployment lifecycle of AI models, from experimentation to production• Collaborate with risk, compliance, and governance teams to ensure all AI deployments adhere to internal policies, regulatory requirements, and ethical AI principles• Lead the response to operational incidents related to deployed AI models, conducting root cause analysis and implementing preventative measures
Qualifications• Education: Bachelor's degree in Computer Science, Engineering, Data Science, Business, or a related field• Experience: Minimum of 8+ years of experience in AI/ML operations, MLOps, DevOps, or a related role with a strong focus on deploying and managing AI/ML solutions in production environments.• Technical Skills:
o Solid understanding of the AI/ML lifecycle, from data preparation and model training to deployment and monitoring.
o Experience with one of the cloud platforms and their AI/ML services
o Proficiency in scripting and
o Familiarity with containerization technologies
o Knowledge of CI/CD pipelines for machine learning models.
o Experience with monitoring tools for AI/ML solutions
o Understanding of data governance, data quality, and data security principles relevant to AI/ML• Strong ability to understand business needs, translate them into technical requirements for AI solutions, and articulate the business value of AI deployments• Excellent communication, interpersonal, and stakeholder management skills• Ability to effectively bridge the gap between technical and business teams• Demonstrated ability to lead initiatives, drive cross-functional projects, and influence outcomes without direct authority• Strong understanding of operational processes and a passion for optimizing them
Corporate Security Responsibility
All activities involving access to Mastercard assets, information, and networks comes with an inherent risk to the organization and, therefore, it is expected that every person working for, or on behalf of, Mastercard is responsible for information security and must:
- Abide by Mastercard's security policies and practices;
- Ensure the confidentiality and integrity of the information being accessed;
- Report any suspected information security violation or breach, and
- Complete all periodic mandatory security trainings in accordance with Mastercard's guidelines.
Similar Jobs at Mastercard
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Leads and develops a cross-functional engineering team building full-stack applications, data pipelines, APIs, and analytics products. Provides technical and product direction, designs complex features, partners on roadmaps, drives innovation and quality, mentors engineers, and coordinates large technical efforts across teams. Contributes directly to implementation while scaling platform solutions and maintaining stakeholder alignment.
Top Skills:
.NetC#JavaPostgresPythonReactReduxSpring BootSQL ServerTypescript
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Leads the architecture, prototyping, engineering, and global deployment of immersive Mastercard Experience Centre technologies. Designs reusable frameworks for AI demonstrations, AR/VR, spatial computing, interactive installations, visualization platforms, IoT environments, and production systems. Provides technical strategy, cross-functional leadership, vendor coordination, operational support models, executive engagement, and mentoring while translating innovation concepts into scalable, reliable customer-facing experiences.
Top Skills:
Agentic AiAPIsArAWSAzureDigital Experience PlatformsDigital TwinsEvent-Driven ArchitecturesGCPGenerative AiHolographic TechnologyInteractive Web ApplicationsIotLarge Language ModelsMobile ApplicationsMulti-Display EnvironmentsProjection TechnologyReal-Time Data PlatformsSensor-Based SystemsSpatial ComputingTouchscreen EcosystemsUnityUnreal EngineVrXr
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Lead Site Reliability Engineer responsible for application reliability, scalability, performance, observability, automation, incident response, capacity planning, and operational risk management. The role develops high-availability solutions, supports production readiness, troubleshoots complex issues, applies cloud and DevOps practices, improves monitoring and resilience, and mentors junior engineers.
Top Skills:
AWSAzureBashCi/CdContainer OrchestrationContainerizationGoGoogle Cloud PlatformLinuxPythonUnix
What you need to know about the Manchester Tech Scene
Home to a £5 billion digital ecosystem, including MediaCity, which consists of major players like the BBC, ITV and Ericsson, Manchester is one of the U.K.'s top digital tech hubs, at the forefront of advancements in film, television and emerging sectors like as e-sports, while also fostering a community of professionals dedicated to pushing creative and technological boundaries.

