SRE Engineer (Site Reliability Engineer) - NetObs Platform
NTT DATA Business Solutions · Bulgaria - Remote · 2026-09-15
Job description
NTT Data Business Solutions is an IT services and solutions provider, primarily engaged with the implementation and maintenance of SAP and other business information systems, IT Infrastructure services, Information Security solutions, Business Development and Project Management. We are working on projects around the world: almost all countries in Europe, the Middle East, the USA and Africa. We are a strategic partner of the biggest SAP Services providers where we are delivering Implementation and Support services. NTT Data Business Solutions is a part of NTT Data Business Solutions Germany.
We are looking for an experienced Site Reliability Engineer (SRE) to join a highly skilled international team supporting a business-critical Network Observability platform. This role combines platform engineering, automation, monitoring, and 24/7 operational support in a secure enterprise environment.
Relocation to Germany is required for this position. The successful candidate will be hired by NTT DATA Business Solutions Germany and must be able to obtain or currently hold a valid German security clearance (Ü2).
Key Responsibilities:
- Operate and maintain Kubernetes-based platforms, including Helm chart management and lifecycle administration.
- Develop and maintain CI/CD pipelines using Jenkins and ArgoCD.
- Manage Infrastructure-as-Code (IaC) solutions to ensure scalable and repeatable environments.
- Create automation scripts using Python, Go, or Bash for provisioning, monitoring, reporting, and operational tasks.
- Configure and support Prometheus, Thanos, and observability solutions.
- Build and maintain Grafana/Perses dashboards and monitoring capabilities.
- Administer Elasticsearch/OpenSearch, Logstash, and Kibana environments.
- Troubleshoot incidents, perform root cause analysis, and participate in Major Incident Management activities.
- Ensure platform security, compliance, and documentation standards are maintained.
- Participate in a 24×7 on-call support model, including weekends and public holidays.
Requirements:
- Proven experience as a Site Reliability Engineer, DevOps Engineer, Platform Engineer, or similar role.
- Strong Linux administration skills, preferably in Kubernetes-based environments.
- Experience with Kubernetes, Helm, Jenkins, ArgoCD, and Git-based workflows.
- Solid scripting skills in Python, Go, or Bash.
- Hands-on experience with Prometheus, Grafana, Elasticsearch, OpenSearch, and related observability tools.
- Good understanding of networking concepts and REST APIs.
- Excellent communication skills in English.
- Willingness to work in a 24/7 support environment.
Nice to Have
- Elastic Certified Engineer certification.
- LPIC Level 2 certification.
- Certified Kubernetes Administrator (CKA).
Mandatory Requirements:
- Citizenship of a country that is a member of both the European Union and NATO.
- Residence in Germany and employment under German labor law.
- Ability to obtain or hold a valid German Security Clearance (Ü2).
What We Offer
- Opportunity to join one of the fastest-growing and most successful IT companies worldwide.
- Work in a collaborative, international, and highly skilled professional environment.
- Exposure to global SAP cloud and digital transformation projects.
- Continuous learning, certification, and career development opportunities.
- Competitive remuneration package and comprehensive social benefits program.
- Flexible working model with remote work options.
- Challenging and rewarding projects with leading international customers.
If you are interested in becoming part of our team, please do not hesitate to send us your resume in English. We thank all interested applicants but will only contact the short-listed ones.