Sr. Site Reliability Engineer (SRE) with Healthcare | Remote | W2 Only
We are seeking a proactive and technically skilled Site Reliability Engineer (SRE) to support production operations of a mission-critical healthcare ecosystem. This role ensures reliability, availability, and operational excellence of healthcare applications, scheduled workloads, and third-party integrations.
The ideal candidate will handle production incident management, Tidal support, troubleshooting healthcare applications, automation scripting, minor enhancements, and inbound/outbound data exchange with vendors. The engineer will collaborate with application, infrastructure, cloud, database, and business teams to ensure uninterrupted healthcare services.
Key Responsibilities
Production Support
-
Provide production support for healthcare applications
-
Monitor environments and respond to incidents per SLA
-
Perform root cause analysis (RCA) and implement preventive actions
-
Participate in on-call rotations and major incident management
-
Ensure high availability and operational stability
Tidal & Batch Processing
-
Monitor scheduled Tidal jobs/workflows
-
Investigate and resolve job failures
-
Restart, rerun, or modify executions per SOPs
-
Analyze dependencies and downstream impacts
-
Coordinate with application/infrastructure teams
-
Recommend automation and self-healing improvements
-
Support batch processing applications
Application & Incident Management
-
ServiceNow incident triage
-
Splunk and Dynatrace analysis
-
Troubleshoot GuidingCare, clinical apps, care management platforms, and integrations
-
Perform configuration validation, performance analysis, and functional verification
Automation & Enhancements
-
Maintain/update scripts (PowerShell, Python, Bash, SQL)
-
Develop automation to reduce manual work
-
Improve monitoring and alerting
-
Support CI/CD operational automation
-
Perform minor code changes, bug fixes, and configuration updates
-
Participate in deployments and validate production fixes
File Transfer & Integrations
-
Monitor file transfers and troubleshoot failures
-
Validate file integrity
-
Manage SFTP/FTPS connections
-
Coordinate with third-party vendors
-
Ensure timely healthcare data delivery
-
Support HIPAA-compliant data exchange
Monitoring & Reliability
-
Monitor applications using enterprise tools
-
Investigate alerts proactively
-
Perform health checks
-
Drive continuous reliability improvements
-
Identify recurring issues and recommend permanent fixes
Documentation
-
Maintain SOPs and knowledge articles
-
Document RCA findings and operational procedures
Required Technical Skills
-
Healthcare application support experience
-
Knowledge of healthcare workflows and HIPAA compliance
-
Healthcare integrations experience
-
Tidal Enterprise Scheduler
-
Job scheduling, batch processing, workflow automation
-
Linux & Windows Server
-
Scripting: PowerShell, Python, Bash/Shell
-
Azure Cloud (preferred)
-
Storage & networking fundamentals
Monitoring Tools (any):
-
Dynatrace, Splunk, Azure Monitor, App Insights
Databases:
-
SQL Server, Oracle
-
Basic query optimization, stored procedures
File Transfer:
-
SFTP, FTPS, file encryption, secure integrations
Version Control:
-
Git / Azure DevOps / GitHub
Preferred Skills
-
ServiceNow
-
CI/CD pipelines
-
REST API troubleshooting
-
Healthcare payer/provider applications
-
AIOps / Agentic AI familiarity
Soft Skills
-
Strong analytical and troubleshooting skills
-
Excellent incident management
-
Ability to work under pressure
-
Strong communication and stakeholder management
-
Customer-focused mindset
-
Continuous improvement attitude
-
Strong documentation skills
-
Ability to work independently and collaboratively
Experience
-
Experience supporting mission-critical healthcare applications (preferred)
-
Agentic AI-driven production operations
-
Auto-remediation and self-healing solutions
-
Azure Logic Apps, Functions, Automation Accounts
-
Healthcare EDI transactions & integrations
-
Experience with GuidingCare or similar platforms