Automated ChatOps in Slack: multi-agent DevOps bot workflow banner
the shed // AGENTIC AI BRIEFING

Every “quick question” in your team channel costs a senior engineer twenty minutes of recovered context. Automated ChatOps makes the channel answer itself, with a hard ceiling on what it is allowed to touch.

See the pattern in action, tap through the tabs below:




chatops-router.yaml

[bottleneck] team-channel-toil

A platform team of nine fields roughly 60 ad hoc requests a week in Slack. Deploy status. Log lookups. Temporary access. Restart this pod. None of it is hard. All of it lands on the same three senior engineers, because they are the only ones who know which dashboard to open.

The cost is not the answer, it is the interrupt. Every question shreds a block of focused work, and the answer disappears into a DM where nobody else can find it. You cannot hire your way out of this, and a wiki does not fix it either, because nobody reads a wiki at 2pm on a Tuesday.

[artifact] chatops-router.yaml

# chatops-router.yaml  --  intent routing, not a chatbot
router:
  model: supervisor
  classify_only: true          # the router never calls a tool itself
  fallback: human_handoff

agents:
  - name: deploy_status
    intents: ["what is deployed", "which sha is in prod", "did my PR ship"]
    tier: 0                    # read-only, answers instantly
    tools: [argocd.get_app, github.get_commit, ci.get_run]

  - name: log_query
    intents: ["why is X 500ing", "show me errors for", "spike in"]
    tier: 0
    tools: [loki.query_range, prom.instant_query]
    limits: { window_max: 6h, rows_max: 200 }

  - name: access_request
    intents: ["I need access to", "grant me", "temporary creds for"]
    tier: 1                    # writes, but gated
    tools: [okta.request_group, vault.issue_lease]
    approval:
      required: true
      approvers: ["@platform-oncall"]
      ttl: 4h                  # every grant expires, no exceptions

  - name: runbook_exec
    intents: ["restart", "drain node", "rotate the pod"]
    tier: 1
    tools: [k8s.rollout_restart, k8s.cordon]
    approval: { required: true, approvers: ["@sre"] }

deny:                          # tier 2, never automated
  - iam.policy_write
  - secrets.rotate
  - db.prod_write
  - billing.*

audit:
  sink: slack_thread + siem
  fields: [requester, intent, agent, tools_called, approver, result]

[trace] 2026-09-16 incident thread

[14:02:11] #platform-help  @jules: anyone know why checkout-api is throwing 500s since the 13:40 deploy?

[14:02:13] shedbot  router  intent=log_query  confidence=0.91  agent=log_query  tier=0
[14:02:14] shedbot  tool=loki.query_range  app=checkout-api  window=30m  rows=147
[14:02:16] shedbot  tool=prom.instant_query  expr=rate(http_5xx[5m])

[14:02:17] shedbot  reply in thread:
  142 of 147 errors are the same trace: UpstreamTimeout on payments-gw.
  Onset 13:41:06, which is 62s after rollout of checkout-api sha 8f21c9d.
  payments-gw p99 went 240ms -> 9.4s at 13:41. checkout-api is the symptom.
  Suggested next step: rollout undo checkout-api, then look at payments-gw.

[14:03:02] @jules: ok roll it back
[14:03:03] shedbot  router  intent=runbook_exec  tier=1  APPROVAL REQUIRED
[14:03:03] shedbot  posts Block Kit card: [ Approve ] [ Deny ]  approvers=@sre
[14:03:41] @kenji (sre) clicked Approve
[14:03:42] shedbot  tool=k8s.rollout_undo  ns=prod  deploy=checkout-api  -> revision 41
[14:04:20] shedbot  5xx rate back to baseline. thread archived to runbook index.

# elapsed human attention: 90 seconds. zero DMs sent.

[checklist] first sprint

1. Export 200 messages from your busiest help channel. Hand-label the top five intents. That label set is your router, and it is the only requirement gathering you need.

2. Ship tier 0 only. Read-only tools, answers in thread, no writes at all for the first two weeks.

3. Add one tier 1 action, the one people ask for most, behind an approval card.

4. Write the deny list before anyone asks for it. IAM, secrets, prod database writes, billing.

5. Measure DM volume to your top three engineers, before and after. That is the metric that justifies the project.

The bottleneck nobody puts on a dashboard

You can measure deploy frequency, MTTR, and change failure rate. You cannot measure the thing that is actually eating your platform team: the steady drip of requests that land in a help channel and get answered by whoever is senior enough to know where to look.

Count them for one week. A team of nine typically finds 50 to 80. Deploy status. “Why is this service slow.” Temporary database read access. Restart a stuck pod. Individually trivial. Collectively, they consume the attention of the three people you most need doing deep work, and the answers vanish into DMs where the next person to ask cannot find them.

This is classic toil in the Google SRE sense: manual, repetitive, automatable, and scaling linearly with the size of your org. The reason it survives is that each request is too small to justify building a tool for. Agents change that math, because you build the routing once and the long tail comes along free.

The outcome: Automated ChatOps

Automated ChatOps is not “put a chatbot in Slack.” It is a specific operational outcome with four properties:

  • Answers land in thread, in public, in under 30 seconds. The channel becomes the documentation, searchable by the next person who asks.
  • Every request is classified before anything runs. A router decides intent, then hands off. No single mega-prompt guessing its way through your infrastructure.
  • Writes are gated, reads are not. Read-only queries answer instantly. Anything that changes state posts an approval card and waits for a human click.
  • There is a deny list, written first. IAM policy, secret rotation, production database writes, and anything touching money never get automated, no matter how convenient it would be.

How it is built

The architecture is a supervisor and a small set of specialists, which is where most teams already land after their first monolithic bot disappoints them.

The router is a cheap, fast model doing one job: map an incoming message to an intent, or to “I do not know, page a human.” Critically, the router holds no tools. It cannot act. That single constraint eliminates most of the ways these systems go wrong.

The specialists each own a narrow domain and a short tool list. A deploy-status agent that can read Argo CD, GitHub, and your CI API. A log agent that can hit Loki and Prometheus with capped time windows and row limits. An access agent wired to Okta and Vault, issuing leases that expire.

The plumbing is unglamorous and matters more than the model choice. Slack Bolt in Socket Mode so you are not exposing an endpoint. Block Kit interactive cards for the approval gate, because a button click gives you an identity and a timestamp for free. An MCP server wrapping your internal APIs so tool definitions live in one place instead of being copy-pasted into four agent configs. LangGraph, CrewAI, or n8n for the orchestration, and the honest answer is that any of them work.

The Block Kit approval card is the load-bearing piece. It turns “the bot did something” into “Kenji approved this at 14:03:41,” which is the difference between a system your security team tolerates and one they shut down.

The three tiers of authority

Write these down before you write any code. The tiers are the product.

  • Tier 0, read. Queries, status, log lookups, config inspection. Runs unattended. Rate limited, scoped, logged.
  • Tier 1, write with approval. Pod restarts, node drains, rollbacks, temporary access grants. The agent proposes the exact call it wants to make, a named human approves, and the grant carries a TTL.
  • Tier 2, never. IAM policy changes, secret rotation, production data mutation, billing. Not gated. Absent from the tool registry entirely, so there is nothing to social-engineer the agent into calling.

The tier 2 list is not paranoia. A campaign disclosed last week showed an attacker driving hundreds of AI agents to compromise 440 PaperCut servers, with one victim going from initial access to domain admin in seven minutes. Agents are fast in both directions. An agent with an IAM write tool is an agent that can grant itself anything, and prompt injection through a pasted log line is a real delivery path.

Where teams get this wrong

Starting with writes. The demo is more impressive and the rollback is more expensive. Ship tier 0 for two weeks first. You will learn which intents actually matter, and you will earn the trust you need for tier 1.

One giant prompt. A single agent with twenty tools and a 4,000-token system prompt degrades unpredictably and is impossible to debug. Route first, then act.

Answering in DMs. Tempting, because it feels tidier. It also destroys the entire benefit. The value is that the answer is public and searchable. Keep it in thread.

No audit sink. If requester, intent, tools called, approver, and result are not landing in your SIEM, you have built a convenience, not a control. That log is what gets this through review.

Your first sprint

  1. Pull 200 messages from your busiest help channel and hand-label the top five intents. This is your requirements doc.
  2. Build the router and exactly one tier 0 agent. Deploy status is usually the highest-volume, lowest-risk starting point.
  3. Instrument the audit sink on day one, not after the security review.
  4. Add one tier 1 action behind a Block Kit approval card once tier 0 is boring.
  5. Track DM volume to your three most-interrupted engineers. A 40 percent drop in a month is a normal result and an easy number to take to your director.

FAQ

How is Automated ChatOps different from an incident response agent?

An incident agent activates when something is already on fire, and its job is triage and escalation. Automated ChatOps runs constantly during normal operations, handling the routine requests that never become incidents. Different trigger, different tool set, different success metric. Most teams get more value from the ChatOps layer first, because the volume is higher and the blast radius is smaller.

Do we need a large model for the router?

No. Intent classification over five to fifteen labels is one of the easiest jobs in the stack, and a small fast model handles it at a fraction of the cost and latency. Save the capable model for the specialist agents that have to reason across log output and correlate timelines. Splitting the tiers this way is also what keeps your monthly spend from surprising you.

What stops someone from talking the bot into doing something dangerous?

Architecture, not prompting. Dangerous tools are not in the registry, so no amount of clever phrasing reaches them. Tier 1 tools require a human click from a named approver, which means the worst case of a successful injection is a rejected approval card and an alert. Never rely on a system prompt as your access control layer.

Want this built for your team?

If your senior engineers are still the routing layer, our DevOps Boot Camp walks through building the tier model and the approval gates end to end. You can also browse all tha-shed courses to find the track that fits where your team is now.