Senior Site Reliability Engineer, AI Infrastructure
PointClickCareRole Overview
PointClickCare is hiring a Senior Site Reliability Engineer, AI Infrastructure. This is a full-time hybrid role, based in Mississauga, Ontario. Part of PointClickCare's Devops hiring, posted 4 weeks ago. Full responsibilities, required qualifications, and the apply link are listed in the description below.
Salary Context
Salary is not disclosed in this posting. Market median for Senior-level Devops roles is $140k-$190k (based on 56 comparable listings). Many employers share specifics during the interview process or after an initial screen.
Resume Keywords to Include
Make sure these keywords appear in your resume to improve ATS scoring
Job description
At PointClickCare our mission is simple: to help providers deliver exceptional care. And that starts with our people. As a leading health tech company that’s founder-led and privately held, we empower our employees to push boundaries, innovate, and shape the future of healthcare.
With the largest long-term and post-acute care dataset and a Marketplace of 400+ integrated partners, our platform serves over 30,000 provider organizations, making a real difference in millions of lives. We also reinvest a significant percentage of our revenue back into research and development, ensuring our employees have the resources to innovate and make a lasting impact. Recognized by Forbes as a top private cloud company and honored as one of Canada’s Most Admired Corporate Cultures, we offer flexibility, growth opportunities, and meaningful work.
At PointClickCare, we empower our people to be the architects of a smarter healthcare future; one that is human-first and accelerated by AI to create meaningful and lasting change. Employees harness AI as a catalyst for creativity, productivity, and thoughtful decision-making. By integrating AI tools into our daily workflows, collaboration is enhanced, outcomes are improved, and every team member has the proficiency to maximize their impact. It all starts with our hiring practices where we uncover AI expertise that complements our mission, and we continue to invest in training and development to nurture innovation throughout the employee journey.
Join us in redefining healthcare — so it doesn’t just survive, it thrives. To learn more about PointClickCare, check out Life at PointClickCare and connect with us on Glassdoor and LinkedIn.
**Travel to Office expectations**
For Remote Roles: If this role is remote, there will be in-office events that will require travel to and from the Mississauga and/or Salt Lake City office. These will include, but not limited to, onboarding, team events, semi-annual and annual team meetings.
For Hybrid Roles: If this role is Hybrid, there will be an expectation to reside within commutable distance to the office/location specified in the job listing. This will include, but not limited to, weekly/bi-weekly/monthly events in the office with your specific team. This is a requirement for this role.
Team Summary:
The AI SRE team is a focused group of SRE engineers dedicated to making PointClickCare's AI and ML platforms reliable, secure, and operationally excellent — from data processing and ML workspaces to model serving and labeling systems. We treat reliability as a product — prioritizing observability, automation, and safe operations so that data scientists and ML engineers can focus on building AI capabilities that improve patient care. You will spend a significant portion of your time hands-on — building automation, designing guardrails, hardening platforms, and leading incident response across cloud environments like Databricks, Azure AI suites. The team collaborates closely with research, platform, data, and security teams across the AI engineering organization.
Job Summary
AI SRE exists to ensure PointClickCare's AI platforms run safely, reliably, and efficiently — protecting patient data while enabling teams to move fast with confidence. We solve complex cross-cutting reliability and security problems through well-designed automation, SLOs, and operational guardrails — so that product and research teams can focus on delivering AI-driven value to clinicians and patients. We own the infrastructure operability of AI data processing, ML workspaces, labeling systems, and model serving — and we drive the infrastructure observability, incident response, compliance controls, and cost optimization that keep those platforms healthy through sound SRE practices and a security-first mindset
Key responsibilities:
Own service level objectives, error budgets, and reliability targets for the infrastructure underpinning cloud-based platforms — ensuring infrastructure observability (metrics, logs, traces), alert quality, and telemetry completeness across platform components and serving endpoints
Design, build, and maintain infrastructure-as-code, operational automation, and change control workflows for AI/ML platforms — with a focus on repeatability, consistency, and toil reduction
Implement and maintain platform security controls — including network segmentation, secrets management, encryption, and data protection safeguards — aligned to compliance requirements and partnering with security teams to respond to emerging risks
Lead incident response and blameless postmortems; validate backup/restore and disaster recovery processes; conduct game days and resiliency testing to harden platform and infrastructure reliability
Mentor engineers, influence design reviews, and collaborate across engineering teams to improve platform resiliency, cost efficiency, capacity planning, and operational standards
Qualification and Skills:
Minimum
5+ years in SRE, platform engineering, or infrastructure roles supporting production cloud environments and mission-critical applications
Strong proficiency with observability — metrics, logging, distributed tracing, SLI/SLO frameworks — and production ownership including incident response, blameless postmortems, and on-call operations
Strong proficiency with Infrastructure as Code (Terraform), GitOps practices, and CI/CD for infrastructure and platform changes
Working proficiency with cloud platform administration — compute, networking, storage, and operating managed data or AI/ML platform services in production (e.g., Databricks, Azure ML, or Kubernetes-hosted infrastructure)
Working proficiency with platform security — network segmentation, secrets management, encryption at rest and in transit, and key management
Strong programming skills for automation, operational tooling, and infrastructure management
Strong communication and documentation skills — able to write runbooks, lead postmortems, influence operational standards across teams, and translate technical complexity for diverse audiences
Preferred
Experience with disaster recovery planning, multi-region patterns, and capacity or cost optimization (FinOps)
Working knowledge of container orchestration (Kubernetes), progressive delivery patterns (blue/green, canary), and data lineage tooling
Working knowledge of container orchestration (Kubernetes), progressive delivery patterns (blue/green, canary), and data lineage tooling
Experience in healthcare, life sciences, or other highly regulated industries with data privacy requirements
About PointClickCare
PointClickCare
pointclickcare.com
68 other open roles at PointClickCare on TryApplyNow.
Frequently Asked Questions
How do I apply for the Senior Site Reliability Engineer, AI Infrastructure position at PointClickCare?
Use the Apply button above to submit your application directly to PointClickCare. Most applications take less than 5 minutes if your resume and contact details are ready, and you'll be routed to the employer's official application system to finish.
Is the Senior Site Reliability Engineer, AI Infrastructure role at PointClickCare remote or in-office?
This is a hybrid role based in Mississauga, Ontario. Expect a mix of in-office and remote days, with the specific cadence set by the hiring manager.
What does a Senior Site Reliability Engineer, AI Infrastructure at PointClickCare earn?
PointClickCare has not disclosed a salary range in this posting. Many employers share specifics later in the interview process; you can also ask during a recruiter screen if compensation transparency is important to you.
When was the Senior Site Reliability Engineer, AI Infrastructure role at PointClickCare posted?
This role was posted on June 23, 2026 (28 days ago). It's still listed as actively hiring; we re-confirm openings against the source system multiple times per day and remove closed roles.
How much experience does the Senior Site Reliability Engineer, AI Infrastructure role at PointClickCare require?
This is a senior-level position. Most senior roles call for 5+ years of directly relevant experience. PointClickCare lists their specific requirements in the description below, so review the must-have qualifications closely before applying.
Similar Jobs
Senior Software QA Engineer, Milpitas, CA, Hybrid
Cisco
Senior Software Engineer, Site Reliability Engineering
Ridgeline International, LLC
Systems Test Engineer, Pipeline and Test Health
Waymo
Senior Product Manager - Network Path
Datadog
Mid-Market Account Executive - San Francisco
Datadog
More Jobs at PointClickCare
View all →Senior Salesforce Architect (CA)
PointClickCare
Senior Vice President & General Manager, Acute & Payer
PointClickCare
VP, Product Delivery & Growth - AI
PointClickCare
(Canada) - Junior Site Reliability Engineer
PointClickCare
Senior Software Engineer- AI Platform(CA)
PointClickCare
AI-powered job search
Get every job scored to your resume
Upload your resume and get jobs ranked, your resume tailored, and employee contacts found automatically.
Get started freeNo credit card to start