Reviewed by Aditya Kumar · Last reviewed 2026-03-24
**Why SLA monitoring**: SLAs are contractual; breaches trigger credits and erode trust. Define SLOs first (e.g., 99.5% success rate, data latency < 30 min). **Architecture**: CloudWatch for metrics—track Lambda/Glue duration, error rates, data freshness (custom metric:...
This hard-level Cloud/Tools question appears frequently in data engineering interviews at companies like Capco. While less common, it tests deeper understanding that distinguishes strong candidates.
This is a senior-level question that tests architectural thinking. Lead with the high-level design, then drill into specifics. Discuss trade-offs explicitly - there is rarely one correct answer. Show awareness of scale, fault tolerance, and operational complexity.
Why SLA monitoring: SLAs are contractual; breaches trigger credits and erode trust. Define SLOs first (e.g., 99.5% success rate, data latency < 30 min). Architecture: CloudWatch for metrics—track Lambda/Glue duration, error rates, data freshness (custom metric: last_successful_run_timestamp). CloudWatch Alarms on threshold breaches (e.g., job duration > 2× baseline, failure count > 0). EventBridge triggers runbooks, PagerDuty, or Step Functions for auto-remediation. Scalability: At 100+ pipelines, custom metrics explode—use namespace + dimension strategy; consider centralized dashboards (Grafana) with templating. Cost: Custom metrics cost $0.30 per metric per month; 50 pipelines × 10 metrics = $150/month—budget for it. For Glue: monitor DPU hours (cost driver) and job bookmarks. Integrate data quality (Great Expectations, dbt tests) that emit pass/fail metrics—quality is part of SLA. Observability: X-Ray for distributed tracing across Lambda + Glue. Store SLAs as code; automate incident response with Step Functions or Lambda so humans handle exceptions, not routine failures.
Pro-Move: Describe implementing an SLA dashboard that shows SLO vs. actual with burn-down—'We track SLO compliance per pipeline and surfaced 3 at-risk pipelines before they breached.' Red Flag: Treating 'pipeline ran' as sufficient—production SLAs require freshness, completeness, and quality metrics.
Some links below are affiliate links. If you buy through them we may earn a small commission at no extra cost to you — it helps keep DataEngPrep free.
According to DataEngPrep.tech, this is one of the most frequently asked Cloud/Tools interview questions, reported at 1 company. DataEngPrep.tech maintains an editor-reviewed database of 1,863 data engineering interview questions across 7 categories.