Site Reliability Engineer · XTB
Apr 2025 – present1 yr 6 mos- Worked with the team to bring separate monitoring systems into one platform for metrics, logs and traces, handling terabytes of telemetry daily.
- Helped define observability standards for 800+ microservices in the cloud and on-premises, including bare metal. Introduced wide events and shared metadata conventions with the team, then worked with other teams to put them into practice.
- Responsible for incident management and on-call for trading services, including SLOs for critical operations, blameless postmortems and runbook automation to shorten recovery time.
- Added JVM metrics and traces to Java services using OpenTelemetry auto-instrumentation, without application code changes.
- Added monitoring for PostgreSQL and Elasticsearch, including slow queries and replication health, plus Kafka consumer lag.
- Built self-service alerting so developers and QA can manage their own alerts and thresholds. Added reports showing each team's alerts, incidents, metric coverage and log usage.
- Built a RAG tool for searching postmortems and internal docs during incidents, a scanner for sensitive data in logs, and an LLM tool that flags incidents in team chat before they are formally reported.
- Plan disaster recovery and run failover drills for trading infrastructure.