IT Craft is looking for an experienced Senior Site Reliability Engineer in Germany.
Key Responsibilities:
- Ensure the reliability, availability, and performance of production systems;
- Define, implement, and continuously improve operational metrics and reliability goals;
- Analyze production incidents, identify root causes, and introduce measures to prevent similar issues in the future;
- Automate repetitive operational processes and minimize manual intervention;
- Build, manage, and continuously improve Kubernetes-based environments and related platform tooling;
- Improve infrastructure scalability and ensure systems can efficiently adapt to changing workloads;
- Strengthen monitoring, alerting, and overall visibility into system health and performance;
- Assess and plan infrastructure capacity requirements for both development and production environments;
- Work closely with software engineering and infrastructure teams to improve system reliability and operational efficiency.
Required Skills:
- Professional experience in SRE, DevOps, platform engineering, infrastructure engineering, or a similar role;
- Strong practical experience managing Kubernetes in production environments;
- Solid understanding of reliability engineering principles, incident response, and system monitoring;
- Experience with infrastructure automation, CI/CD, and automated deployment processes;
- Good knowledge of compute, storage, networking, and other core infrastructure concepts;
- Experience with capacity planning, system scaling, and performance optimization;
- Willingness to explore and adopt new engineering tools, technologies, and automation practices;
- English proficiency at Upper-intermediate level or above; German would be an advantage.