TL;DR: agency-agents is the largest curated collection of AI agent personas on GitHub, at 149,312 stars as of September 1, 2026. It ships more than 300 agent definition files across roughly 20 category folders and installs them into Claude Code, Cursor, Codex, Gemini CLI, OpenCode, Windsurf, Aider, and a dozen other tools with one command. What you get is framing, not capability. The personas change how your agent approaches a task; they do not give it facts it does not have, and they do not survive past the session.
This is a deep dive on one tool from our roundup of five open source AI agent tools worth installing in 2026.
Your coding agent has one personality: helpful generalist. Ask it to review an authentication flow and you get sensible, forgettable advice. Ask a security reviewer with a threat model and a checklist, and you get findings.
That gap is what agency-agents fills. It started as a Reddit thread about agent specialization and grew into a roster large enough that installing all of it breaks at least one popular agent runtime. This is what is actually inside, how to install the parts you want, and the two things a persona library cannot do for you.
What you are actually installing
Each agent is a markdown file. Not a one-line system prompt, and not a plugin with code. A file contains an identity and personality, a core mission, a working process, technical deliverables with examples, and success metrics.

Counting the markdown files in the repository tree on September 1, 2026, there are 312 across 20 top-level category folders, or 306 once you exclude the examples/ directory. The distribution is lopsided toward building things:
| Division | Agent files |
|---|---|
| engineering | 59 |
| specialized | 58 |
| marketing | 36 |
| game-development | 21 |
| integrations | 18 |
| strategy | 16 |
| gis | 13 |
| security | 12 |
| design | 10 |
| sales | 9 |
| testing | 9 |
| paid-media | 7 |
| project-management | 7 |
| academic | 6 |
| spatial-computing | 6 |
| support | 6 |
| finance | 5 |
| product | 5 |
| healthcare | 3 |
Worth noting the repository README still advertises “230+ agents.” That line is stale. The tree has grown past 300.
The engineering division is where most developers will start, and the specificity is higher than you would expect from a prompt collection. Alongside the obvious Frontend Developer and Backend Architect entries there is a Network Engineer scoped to Cisco IOS-XE, Juniper Junos, and Palo Alto PAN-OS, an Embedded Firmware Engineer for ESP32, STM32, and Nordic targets, an Incident Response Commander for post-mortems and on-call readiness, and a Codebase Onboarding Engineer written to explore a repository read-only and state facts about it rather than propose changes.
That last one is a good example of the pattern working. A read-only persona with an explicit no-edit mandate is a genuinely different tool from your default agent, and it costs nothing but a file.
Installing without breaking your setup
The repository ships conversion and install scripts. The interactive path detects what you have installed and asks what you want:
git clone https://github.com/msitarzewski/agency-agents.git
cd agency-agents
./scripts/install.sh
Targeting a specific tool and a subset of divisions is the saner default:
# everything, into Claude Code
./scripts/install.sh --tool claude-code
# two divisions only
./scripts/install.sh --tool claude-code --division engineering,security
# named agents only
./scripts/install.sh --tool cursor --agent frontend-developer,ui-designer
# see what exists before committing
./scripts/install.sh --list teams
./scripts/install.sh --tool opencode --division engineering --dry-run
Supported targets include Claude Code, Cursor, Codex, Gemini CLI, OpenCode, GitHub Copilot, Windsurf, Aider, Kimi Code, Hermes, Antigravity, Osaurus, and Mistral Vibe. There is also a native desktop app at agencyagents.app for macOS, Linux, and Windows that browses the roster, installs with a click, and auto-updates, plus a Homebrew cask.
Read this before you install everything. OpenCode’s runtime currently registers only about 119 agents and silently drops the rest, which the repository documents against an upstream bug. Installing a subset with --division keeps you under the limit, and the installer warns you when a selection would exceed it. Silent truncation is the worst failure mode there is, because the agent you wanted is missing and nothing tells you.
Even outside that specific bug, installing 300 personas is a bad idea. A roster you cannot remember is a roster you will not use. Install the two divisions you work in, read four or five of the files, and delete the ones that do not match how your team actually operates.
What a persona changes, and what it does not
A persona is a frame. It sets what the agent looks for, what format it produces, and what it considers done. That is worth more than it sounds, because a large part of bad agent output is not the model being wrong, it is the model optimizing for the wrong shape of answer.
What a persona cannot do is supply information the model does not have. This is the limit people discover a week in, and it shows up hardest around APIs.
Load the Backend Architect and ask it to write a client for your internal billing service. It will produce clean, idiomatic code against a response shape it invented. It has no way to know your service returns a 409 with a different error envelope when an idempotency key repeats, or that the pagination cursor is opaque rather than an offset. The persona made the code better organized. It did not make it correct.
The same limit applies to the testing division. A QA persona writes thorough tests against its own mental model of the API, so the tests pass and prove nothing. We covered why that specific loop is dangerous in testing non-deterministic AI agents, and what breaks when the real shape moves in what happens when API changes break AI agents.
The fix is to give the agent the contract instead of hoping the persona compensates. If your API is designed in Apidog, the OpenAPI specification is the source of truth: real schemas, real status codes, real error envelopes. The agent reads them rather than reconstructing them from call sites. Mocks generate from that same spec, including the error branches a persona would never think to fake, and the test suite fails loudly in CI when reality and spec diverge.

That pairing is the useful mental model. The persona decides how the agent works. The spec decides what is true. Related reading: using your OpenAPI spec as agent tools and designing API tool schemas for agents. If you want to wire a live spec into your agent’s context, download Apidog and point it at your existing project.
The second limit: a persona is not a team
Here is the thing the roster metaphor promises and does not deliver.
The pitch is a complete agency at your fingertips. In practice, you activate one persona in one session, it does a task, and the session ends. Tomorrow you type the activation again. There is no team, because there is nothing persistent: no assignment, no record of what the security reviewer found last Tuesday, and no way for a colleague to see any of it. Nine of the divisions in this repository describe roles that only make sense inside an organization, and the tool that installs them has no concept of one.
If you want the roster idea to survive past a single terminal session, the missing layer is work management. Sharkly is built for exactly that, and the mapping onto agency-agents is close enough to be worth spelling out.

- An Agent in Sharkly is a saved configuration, not a prompt you retype. It holds instructions, runtime, skills, repositories, and environment. A persona you tuned once gets reused, which is the durable version of what an
.mdfile in~/.claude/agents/is reaching for. - A Crew is a leader agent plus other agents and people. This is the division concept with actual mechanics. Crews run leader-first: the leader reads the task context, decides which members to pull in, and combines their results in one place, rather than every specialist starting at once and racing.
- Work is assigned, not invoked. You give a task to an Agent the way you give it to a teammate. It lives in a space, a project, and a sprint, with Jira sync if your team already works there.
- Execution is bring your own. You connect a Computer, which can be your laptop, a server, or a container, and Sharkly uses the Runtime already installed on it. Your Claude Code or Codex subscription does the work; nothing resells you tokens.
- The output is reviewable. Progress, tool calls, and results stream back to the task, and the agent’s output lands as comments you reply to. A task in Backlog does not start a run, so you can prepare work before anything executes.
Put plainly: agency-agents gives you the job descriptions, and Sharkly gives you the place those jobs get assigned, run, and reviewed. Use the repository as a library of role definitions to seed real Agents rather than as a folder you install and forget.
Writing your own, using theirs as the template
The most durable thing this repository gives you is a format. Once you have read a few files, writing a persona for your own stack takes about fifteen minutes, and a house-specific one beats a generic one every time.
The structure that repeats across the good files:
- Identity and voice. Who this agent is and how it talks. Sounds cosmetic, and it is the part that stops the agent drifting back into generic assistant mode halfway through a long task.
- Core mission. One sentence on what success means. The Codebase Onboarding Engineer file is a clean example: explore read-only, trace code paths, state facts about structure and behavior, propose nothing.
- Working process. Ordered steps the agent follows. This is the section worth stealing. A review persona with a seven-step process produces consistent output across sessions; one without it produces whatever the model felt like that day.
- Deliverables. The concrete artifacts, with format. Say “a markdown table of findings with severity, file path, and line number” rather than “a report.”
- Success metrics. How the agent knows it is done, which doubles as the thing you check its work against.
A house version for API work might look like this:
# API Contract Reviewer
## Mission
Verify that new or changed endpoints match the OpenAPI spec in this
repository before they reach review. Report mismatches. Do not edit code.
## Process
1. Read the spec for every endpoint touched by the current diff.
2. For each one, compare the handler against the spec: status codes,
response schema, error envelope, required headers, pagination style.
3. Run the contract tests. Record failures verbatim.
4. Check that new endpoints were added to the spec, not just to the router.
5. Flag any response field present in code and absent from the spec.
## Deliverables
A table: endpoint, method, mismatch type, spec line, code line, severity.
No prose summary. No suggested fixes unless asked.
## Done when
Every endpoint in the diff appears in the table with a verdict, and the
contract test output is included as evidence.
That file is short, and it does more for a backend team than any of the 59 engineering personas shipped in the repo, because it names your spec, your tests, and your definition of done. Step three is the part that matters: the persona is instructed to produce evidence rather than an opinion, which is the difference between a review you can act on and a paragraph of reassurance.
The same trick works for incident response, migration work, dependency upgrades, and onboarding. Take the file structure, keep the process discipline, replace the generic content with yours. When agent output has to survive a handoff to a human or another agent, the format contract does most of the work, which is the point we made in agent handoff and context passing.
Are the star counts meaningful here?
149,312 stars is a lot, and it deserves a caveat. Persona repositories collect stars faster than almost any other category because they are easy to understand, easy to share, and cost nothing to try. A star means somebody thought this was a good idea, not that they still use it.
What makes this repo credible is the shape of the work rather than the number. It has real contribution history since October 2025, it is MIT licensed, it documents its own limits including the OpenCode registration cap, and the agent files carry specific domain content rather than generic role fluff. Compare that to the many “awesome prompts” repos that peaked and stopped.
Judge it by opening three files in a division you know well. If the Frontend Developer file says things a good frontend developer would say, the rest is probably fine. If it reads like a job posting, skip the repo.
How to actually get value from it
A workflow that holds up:
- Install one division. Pick the one matching your daily work. Engineering for most readers.
- Read the files. Four or five, start to finish. You are looking for the process sections, which are the part worth keeping.
- Edit them. Add your stack, your conventions, your definition of done. A persona describing a generic React shop is worth less than the same file with your testing rules in it.
- Give the edited personas real inputs. A security reviewer with your OpenAPI spec finds contract problems. The same reviewer without it produces a checklist.
- Promote the ones that stick. Any persona you reach for twice a week should become a saved Agent in a work management system rather than a file you copy between machines.
That last step is where teams either scale this or quietly abandon it.
FAQ
Does agency-agents work with Cursor and Codex, or only Claude Code? All of them. The convert.sh script generates integration files per tool and install.sh --tool targets a specific one. Claude Code, Cursor, Codex, Gemini CLI, OpenCode, Copilot, Windsurf, Aider, Kimi Code, and several others are supported. If you are comparing agent clients for API work, see our look at API clients in Cursor and Copilot.
Should I install all 300 agents? No. On OpenCode you cannot, because the runtime registers roughly 119 and drops the rest silently. On other tools you technically can, and you will never use most of them. Install by division.
Do personas make the agent smarter? They make it better directed. Accuracy on questions of fact, including what your API returns, does not move. That needs a real specification, which is the argument in do you still need an API tool in the age of AI agents.
Is it safe to run the install script? It writes agent definition files into your tool’s config directories. That is its purpose. It is MIT licensed with the full source available, and --dry-run shows you what it would do before it does it. Run the dry run first. The general principle for what you let agents touch is in AI agent guardrails.
What is the difference between this and an agent framework? If you want the personas running several at a time rather than one per session, that is Orca. A framework like Strands or AgentKit gives you runtime orchestration in code. agency-agents gives you markdown files that change how an existing agent behaves. Different layer, no overlap.
Wrapping up
agency-agents is the best version of a specific idea: your agent does better work when it knows what role it is playing. Install a division, read the files, edit them into something that matches your team, and you will get real value for about twenty minutes of setup.
Then be honest about the two things it does not solve. The personas do not know what your API returns, which is what Apidog is for, and they do not persist into anything a team can see or review, which is what Sharkly is for. A roster is a good start. It is not an organization, and it is not a source of truth.



