CrewAI gives you role-based agents in YAML and a Flow to keep them on rails. Here is how to turn that into incident write-ups, drift reports, and CVE sweeps that land as files, not chat.
See the workflow in action, tap through the tabs below:
$ terminal // Python 3.10 to 3.13, uv required
uv tool install crewai crewai create flow ops-flow cd ops_flow # add SERPER_API_KEY and your model key to .env crewai install crewai run
$ src/ops_flow/crews/triage_crew/config/agents.yaml
log_reader:
role: > {service} Log Triage Analyst
goal: > Find the first error, the blast radius, and the likely cause
backstory: > SRE who reads logs bottom-up and never guesses without a line number.
incident_writer:
role: > Incident Scribe
goal: > Turn the analyst notes into a timeline and a draft postmortem
backstory: > Writes blameless, specific, and short. Cites log lines.
$ src/ops_flow/main.py
class TriageState(BaseModel):
service: str = ""
report: str = ""
class OpsFlow(Flow[TriageState]):
@start()
def pick_service(self):
self.state.service = "payments-api"
@listen(pick_service)
def run_triage(self):
out = TriageCrew().crew().kickoff(inputs={"service": self.state.service})
self.state.report = out.raw
YAML keys must match method names. log_reader in agents.yaml needs a def log_reader(self) decorated with @agent. Mismatch means a confusing KeyError.
Give agents read-only tools first. A crew with a shell tool and a vague goal will happily run commands you did not mean. Start with search and file readers.
Every agent call is tokens. A three-agent crew with verbose reasoning can cost 10x a single prompt. Cap iterations and log usage.
What CrewAI is, in one paragraph
CrewAI is an open-source Python framework for multi-agent systems. You describe agents by role, goal, and backstory in YAML, give them tasks with an expected output, and group them into a Crew that runs sequentially or hierarchically. Sitting above that is a Flow: an event-driven, stateful orchestrator with @start, @listen, and router steps that decides when each crew runs and what it gets. Crews do the thinking, Flows keep them on a leash. For ops people that split is the reason to pick CrewAI over gluing prompts together: you get autonomy where it helps (reading logs, comparing configs) and determinism where it matters (what runs, in what order, with what state). If you want the broader landscape, we compared frameworks in our Make AI Agents post.
Quick setup
CrewAI needs Python 3.10 through 3.13 and uses uv for everything. Install the CLI with uv tool install crewai, then scaffold a Flow project with crewai create flow ops-flow (the folder becomes ops_flow). You get src/ops_flow/main.py for the Flow and a starter crew under crews/ with config/agents.yaml and config/tasks.yaml. Put your model provider key and a SERPER_API_KEY (for web search) in .env, run crewai install, then crewai run. If you only want a crew with no Flow, crewai create crew name gives you the smaller scaffold. Coding agents can pick up the framework's conventions with npx skills add crewaiinc/skills, which installs CrewAI skills for Claude Code and similar tools.
The mindset: agents are junior hires with a job description
The YAML is not decoration. An agent with the role "Log Triage Analyst," a goal that names the deliverable, and a backstory that says "never guesses without a line number" behaves measurably better than "helpful assistant." Write the role the way you would write a job posting for a junior SRE. Then write the task's expected_output the way you would write acceptance criteria: sections, length, format, what not to include. Finally, keep the Flow boring. State in a Pydantic model, one crew per step, a router only where a real decision exists. If you find yourself giving an agent a shell tool "so it can figure it out," stop. That is the moment to add a step to the Flow instead.
7 crews worth building this week
1. Incident triage crew
Two agents: a log reader with a file-read tool pointed at an exported log bundle, and a scribe. Task one: find the first error, the blast radius, and the likely cause with line numbers. Task two: a timeline and a blameless draft postmortem, output_file: output/incident.md. Kick it off from a Flow that takes the service name as state. You get a draft in two minutes that a human edits, not a chat you have to copy out of.
2. Config drift reporter
Give one agent the rendered Terraform plan and another the live inventory export (both as files). Expected output: a table of resources where code and reality disagree, sorted by security impact, with "open to 0.0.0.0/0" flagged first. Run it nightly from a cron that calls crewai run. This pairs nicely with the deterministic detectors we built in Automated Drift Detection; the crew writes the narrative, the detector supplies the facts.
3. CVE sweep crew
A researcher with SerperDevTool and a dependency-file reader. Task: read every lockfile in a repo list, cross-reference against advisories published this month, and produce a table of repo, package, version, CVE, severity, fix version. Add a second task that drafts the upgrade PR description. Read-only tools only. Nothing here should touch git.
4. Runbook gap finder
Feed the crew your alert rules export and your runbook directory. Expected output: every alert without a linked runbook, plus a one-paragraph stub for each written in the style of the existing runbooks. This is the highest ratio of "boring but valuable" to effort on the list.
5. Change-review crew with a human gate
Use a Flow router. If the crew rates a change as high risk, route to a review step that waits for a person before anything else happens; otherwise write the summary to disk and stop. CrewAI Flows support human-in-the-loop patterns for exactly this. The rule: the model can recommend, the Flow decides, the human approves writes.
6. Access-review crew
Input: an IAM policy export and an HR roster CSV (generic, no real data in your test runs). Tasks: list identities with admin-equivalent rights, flag any that are not on the roster or whose role does not justify the rights, and draft the quarterly access-review memo. Great candidate for Process.hierarchical with a manager agent delegating the two analyses.
7. On-call handoff crew
At shift end, run a crew over the last 12 hours of ticket and alert exports. Output: what fired, what is still open, what the next person needs to know, and one line on anything that looks like it will fire again. Save it to output/handoff.md and post it yourself. Your future self at 3 AM will thank you.
Safety and gotchas
Names must line up. The YAML key log_reader has to match a method named log_reader decorated with @agent on your @CrewBase class, and the same goes for tasks. Most first-run failures are this.
Tools are the attack surface. An agent with a shell or HTTP-write tool plus a tool result containing "ignore your instructions" is a prompt injection waiting to happen. Start every crew with read-only tools (file read, search). Add write capability only inside a Flow step a human gates.
Autonomy is not free. A verbose three-agent crew can burn ten times the tokens of a single well-written prompt. Set a max iteration count on agents, keep expected_output tight, and watch the first few runs with verbose=True before you cron anything.
Never test on real data. Build your crews against synthetic log bundles and fake rosters first. The example data in this post is invented for a reason.
Output files are the contract. Treat output/*.md as the API of your crew. If downstream automation parses it, pin the format in expected_output and add a validation step in the Flow.
Usage and cost tips
CrewAI itself is open source and free. What you pay for is the model behind it and, optionally, hosting. Sequential process with two or three agents is the sweet spot; hierarchical adds a manager agent and more calls, so use it only when delegation is the point (the access-review crew above). Pick a cheaper, faster model for reader-type agents and reserve your strongest model for the scribe or reviewer. Cache inputs where you can: a crew that re-reads a 40 MB log on every run is wasting money, so pre-filter with grep before kickoff. When a crew works locally, crewai login and crewai deploy create push it to CrewAI AMP from a GitHub repo, with crewai deploy status and crewai deploy logs for the day two stuff. Run plot() on your Flow first: the generated diagram is the fastest way to spot a step that should not exist.
FAQ
Do I need a Flow, or is a Crew enough?
A Crew alone is fine for one-shot jobs you run by hand. The moment you need state between steps, a decision (route on severity), a human gate, or a schedule, use a Flow. CrewAI's docs call Flows the recommended structure for production apps.
Can CrewAI agents use my existing tools and APIs?
Yes. crewai_tools ships file, search, and web tools, and you can wrap any Python function or API as a custom tool. Keep production tools read-only until a Flow step explicitly gates writes.
Which Python version does CrewAI support?
Python 3.10 up to but not including 3.14, managed with uv. Install the CLI with uv tool install crewai and verify with uv tool list.
Build one crew today
Scaffold the Flow, write two agents in YAML, point them at a synthetic log bundle, and get output/incident.md on disk. That is a one-evening project, and it teaches you more about agent design than a week of reading. If you want structured practice turning this into a career skill, see our courses.
