Senior Site Reliability Engineer, Production Engineering Lisbon, Portugal

Detalhes da Vaga

Senior Site Reliability Engineer, Production EngineeringLisbon, Portugal
Please note that we have a hybrid approach to work and would like to find someone who can come into our offices in Lagoas Park once a week.Who We AreCisco ThousandEyes is a leading Digital Experience Assurance platform that empowers organizations to deliver seamless digital experiences across every network—even those beyond their ownership. Leveraging AI and an unparalleled set of cloud, internet, and enterprise network telemetry data, ThousandEyes enables IT teams to proactively detect, diagnose, and resolve issues before they impact end-user experiences.
About The RoleWe are seeking a skilled Senior Site Reliability Engineer (SRE) in Production Engineering with a strong background in SaaS and operations. You will design and manage large-scale, highly available distributed systems in the cloud, collaborating directly with application development teams to enhance the reliability, performance, and security of our platform.
Key ResponsibilitiesIdentify and provide solutions to common obstacles hindering operational excellence across engineering teams.Partner with application developers using cloud-native tools to address novel challenges around scale, performance, and reliability.Generalize and standardize solutions and processes to enable repeated success across our microservice-based multi-region platform.Play a key role in the ThousandEyes platform by leveraging scale testing, additional environments, and working with application teams to improve system reliability.Use cloud-native observability and reliability tools such as Prometheus, Istio, and ArgoCD.Manage a rapidly growing infrastructure capable of handling substantial daily data volumes, emphasizing operations/infrastructure/everything as code.What You'll DoCollaborate with software engineers to ensure architecture and services are optimized for availability, latency, and performance.Design and implement scalable operations tooling to support platform growth and scaling across multiple regions.Design, deploy, and maintain AWS cloud-native services that are elastic and resilient to failure.Participate in and improve our 24x7 incident response and on-call rotation.Use and expand our existing CNCF solutions like Kubernetes, Service Mesh, Prometheus, OpenTelemetry, and ArgoCD to increase platform reliability.Automate production operations to provide guardrails and continuous platform operation.Develop automation solutions for scalable service and platform operations, including deployment, scale testing, graceful failure, and chaos testing.Stay updated on industry best practices for scalability and reliability to improve the scalability of the ThousandEyes platform.Required QualificationsExpert-level knowledge of Kubernetes and its ecosystem.Proficiency in software development with languages such as Python or Go.In-depth knowledge of cloud providers, preferably AWS.Proven ability to build and implement scalable and well-tested solutions.Strong understanding of Unix/Linux systems, including kernel, system libraries, file systems, and client-server protocols.Knowledge of Site Reliability principles: Incident Response, Change Management, Distributed Systems, Deployment Strategies, and SLOs.Excellent communication and documentation skills.Strong sense of ownership, drive, and attention to detail.Preferred QualificationsFamiliarity with best practices for operating a large-scale, highly available enterprise platform.5+ years of experience in a related role.Cisco values the perspectives and skills that emerge from employees with diverse backgrounds. We believe that everyone has something to offer and that diverse teams are better equipped to solve problems, innovate, and create a positive impact. We encourage you to apply even if you do not believe you meet every single qualification.
Cisco is an Affirmative Action and Equal Opportunity Employer and all qualified applicants will receive consideration for employment without regard to race, color, religion, gender, sexual orientation, national origin, genetic information, age, disability, veteran status, or any other legally protected basis.

#J-18808-Ljbffr


Salário Nominal: A acordar

Fonte: Jobleads

Função de trabalho:

Requisitos

Técnico De Manutenção Eletromec Nico - Sintra

TÉCNICO DE MANUTENÇÃO ELETROMECÂNICO - SINTRA 2024-10-25 Sintra, Lisboa Manutenção, Instalação, Reparação Técnica Ref: 105-004620-1 ANÚNCIO DE VAGA: TÉCNICO ...


Adecco Prestação De Serviços, Lda - Lisboa

Publicado 11 days ago

Diretor De Obra - Lisboa (M/F)

A Mota-Engil ATIV é uma nova marca no universo do Grupo Mota-Engil que alavanca o reconhecido conhecimento e experiência da Manvia e da Vibeiras, promovendo ...


Mota-Engil - Lisboa

Publicado 11 days ago

Head Of Mechanical Engineering

At Bloq.it, we've created the world's leading smart locker solution. Solving online deliveries by enabling everyone to participate easily, reducing delivery ...


Bloq.It - Lisboa

Publicado 11 days ago

Head Of Maintenance Workshop And Mobile Equipment

Head of Maintenance Workshop and Mobile EquipmentPosition to be based in Luanda, AngolaKey purpose:The Head of the Workshop will report to the Head of the ma...


Boardroom Appointments - Lisboa

Publicado 11 days ago

Built at: 2024-12-24T18:15:24.714Z