Site Reliability Engineer (SRE) – Application Support
Diversified Services Network, Inc. (DSN) is seeking a full-time Site Reliability Engineer (SRE) – Application Support to join our team in their choice of our Chicago, IL or Peoria, IL office locations. We offer a full hybrid work model, requiring 2 days of onsite work per week, benefits, PTO, 401k, and more! If you're looking to grow your career within an extremely reputable, stable Fortune 100 company - let's talk!
Position Overview
We are seeking a Site Reliability Engineer (SRE) – Application Support to triage, investigate, and resolve support tickets for critical business applications, ensuring availability and performance are maintained against defined Service Level Agreements. This role combines hands-on application support with automation and monitoring work, and includes periodic off-shift and weekend support coverage.
Key Responsibilities
-
Triage and resolve support tickets in accordance with defined Service Level Agreements (SLAs).
-
Identify and resolve critical application and technical problems, including responding to off-shift and weekend support calls.
-
Investigate tickets and perform statistical analysis and testing of application failures.
-
Proactively identify and manage issue resolution, documenting the process and providing follow-up to requestors.
-
Monitor application availability and performance on an ongoing basis.
-
Develop scripts and automation tools to better detect and correct application issues.
-
Build monitoring and alerting capabilities to proactively surface application issues.
-
Collaborate cross-functionally, including occasional coordination with external partner/dealer communications.
-
Perform additional duties as assigned.
Requirements
Education & Experience
-
Bachelor's or Master's degree with 2–4 years of experience in Application Support for cloud-based applications; OR
-
No degree with a relevant technical certification and 6+ years of comparable experience.
Required Technical Skills
-
Expertise supporting, querying, and reporting on relational databases, preferably Snowflake.
-
Experience with AWS services: EC2, S3, VPC, Route 53, RDS, CloudFormation, DynamoDB, Lambda, CloudWatch, IAM, Certificate Manager, ELB, EBS, ECS, CloudFront/WAF, SQS, SNS, and SES.
-
Experience troubleshooting issues related to UI, API, and data flow.
-
Expertise with high-availability architecture.
-
Experience with monitoring tools (e.g., ThousandEyes, AppDynamics, Grafana).
-
Background in data management, data engineering, or data operations; familiarity with ADO pipeline framework a plus.
-
Experience with Python or a related scripting language for automation.
-
Experience in a Site Reliability Engineering (SRE) capacity, improving application availability and performance.
-
Working knowledge of Python, PowerShell, SQL, and JSON.
-
Working knowledge of autoscaling deployment concepts (hands-on configuration not required).
Benefits
-
401(k)
-
Dental insurance
-
Vision Insurance
-
Disability insurance
-
Employee assistance program
-
Health insurance
-
Health savings account
-
Life insurance
-
Paid time off
-
Paid Holidays
Please follow the link to our website for a list of job openings in Engineering, IT, Project Management, and more! https://www.dsnworldwide.com
Salary expectations: 95,000-100,000 per annual