Architecture and ownership
List critical services, owners, dependencies, public entry points, deployment paths, and rollback steps. Reliability starts with knowing what exists and who can change it.
Observability and backups
Check alerts, dashboards, log retention, backup schedules, restore tests, and access to production evidence. A backup that nobody has restored is an assumption, not a plan.
Cost and security drift
Reliability reviews should include waste and exposure. Unused resources, stale access, and unclear network rules are operational risks as well as cost or security problems.

