Recognizable symptoms
- The platform works, but nobody is confident operating it under incident pressure.
- Security, cost, and rollback decisions are spread across tickets, memory, and old deploy scripts.
- Monitoring exists, but alerts do not clearly explain user impact or ownership.
- The next production launch depends on senior people being online at the right moment.