Start here
Where SRE and platform engineering need an explicit handoff
Define where platform responsibilities meet service reliability.
Read the articleWork through infrastructure, capacity, deployment and ownership decisions behind dependable services.
Start here
Define where platform responsibilities meet service reliability.
Read the articleUnderstand the platform
Examine process packaging, orchestration and application responsibilities.
Read the articlePlan maintenance
Work through a node drain and the capacity checks around it.
Read the guideControl a rollout
Review traffic selection, promotion criteria and rollback compatibility.
Read the guideNews ·
Codex Cloud reuses prepared repositories and tools while keeping task files separate. Learn what an environment update changes and what it leaves running.
News ·
Google’s legacy agents have entered a migration grace period. Compare replacements and check the queries that must survive the switch.
News ·
Google’s Storage Intelligence update connects usage findings with bulk changes. Dry runs help check the selection before objects are changed.
Guide · 11 min read
Learn what CodeRabbit does, set up AI code review, understand pricing and privacy, and decide when it helps your team or adds more work.
Guide · 8 min read
Herdr keeps coding agents visible and their terminals persistent. Here is how to evaluate its status signals, recovery behavior, and place in SRE work.
Explainer · 4 min read
When inference slows without obvious application errors, test physical constraints alongside queueing, workload changes, and software regressions.
Guide · 4 min read
Turn reliability policy into a release decision with fresh evidence, limited exceptions, and a clear path to resuming normal changes.