The library

Everything we index — ranked by what works, never by stars.

untested
Generate and rank candidate plansworkflowProductOpsL2
judge-panel · Assembling diverse expert perspectives in a structured panel verdict
untested
Deep multi-agent flow verificationworkflowEngineeringL3
deepcheck · Deep verification of code correctness with multi-lens checking
untested
Stress-test design through lensesworkflowProductEngineeringL2
design-review · Systematically reviewing design artifacts across consistency and usability
untested
Audit pages and build missing featuresworkflowEngineeringProductL3
path-warden-completeness · Verifying task completion paths are complete and well-formed
untested
Run test-driven development cycleworkflowEngineeringL3
tdd-cycle · Testing TDD workflows when you need automated red/green/mutation gates without interactive review stops.
untested
Review pull requests with Opus specialistsworkflowEngineeringL3
athena-pr-reviewer-workflow · Auditing PRs when you need parallel specialist analysis (comment, test, error, type, code, simplify, requirements) with batched verification.
untested
Audit agent framework primitivesworkflowEngineeringL3
fs-harness-research · Researching filesystem abstractions when comparing Flue, Hare, cloudflare-agents, Mastra, and Pi implementations.
untested
Audit codebase for tool vulnerabilitiesworkflowEngineeringL3
codebase-audit · Auditing codebases when you need whole-repo cross-slice analysis (orphans, duplication, dead code, architecture drift, HARD-RULE compliance).
untested
Expand multilingual fact databaseworkflowDataL3
hiraia-expand-wave · Building fact banks when you need medium-grained concept beats expanded into distinct trilingual (Tagalog/English/Bisaya) child-grade facts.
untested
Audit inspector form usabilityworkflowProductL3
inspector-usability-audit · Auditing UI form builders when you need full data-flow verification (form→state→serialization) against upstream schema docs.
untested
QA rendered content for extraction errorsworkflowOpsL3
content-rendering-qa · QA text content when you need to detect HTML/entities, page numbers, OCR gibberish, dropped words, truncation in rendered fields.
untested
Implement GitHub issues end-to-endworkflowEngineeringL3
liftoff-workflow · Project execution when you need parallel Astronauts to implement task batches with FC verification and worktree commits.
untested
Compare specialist agents head-to-headworkflowEngineeringL3
«family»-agent-headtohead · Benchmarking agent variants when you need head-to-head evaluation on a fixed task set.
untested
Generate multilingual course vocabularyworkflowMarketingL2
lang-content-generate · Content localization when translating a spec into multiple languages (often with cultural/regional variants).
untested
Refactor code without behavior changeworkflowEngineeringL3
refactor · Improving code quality when tests are passing and you want to safely restructure without behavior change.
untested
Optimize Google Ads campaigns automaticallyworkflowMarketingL3
optimization-loop · Improving performance when you need goal-driven iteration (profile → hypothesis → implement → measure).
untested
Plan proposals with adversarial reviewworkflowOpsL3
wf-plan · Designing workflows when you need to decompose a goal into agent-ready phase briefs and schemas.
untested
Design in-silico apoptosis research programworkflowEngineeringL4
deep-insilico-program-design · Algorithm/architecture design when you need iterative sketches stress-tested by adversaries before coding.
untested
Batch improve and refresh skillsworkflowOpsL2
skill-improver-batch · Skill refinement when you have many skills needing updates based on failure logs or user feedback.
untested
Evaluate workflow runtime systemsworkflowEngineeringL3
evaluate-external-workflow · Validating imported workflows when you need to verify they work on your task set.
untested
Execute complete feature lifecycleworkflowEngineeringL3
feature · Feature delivery when you need end-to-end orchestration (design → implement → test → deploy).
untested
Analyze source files for needed testsworkflowEngineeringL2
test-analysis · Debugging tests when you need to distinguish flakiness from real failures and identify patterns.
untested
Harden trigger evaluation with adversarial queriesworkflowEngineeringL3
trigger-eval-harden · Automating testing when you need to harden Workflow triggers before production deployment.
untested
Auto-generate unit tests with coverageworkflowEngineeringL2
unit-test-gen · Test coverage when you have code without tests and need to bootstrap test suite.
untested
Run load tests with ramping VUworkflowEngineeringL2
load-test · Finding capacity limits through stepped load with automatic breaking-point detection.
untested
Validate dynamic roster skill invariantsworkflowEngineeringL2
fixture-dynamic-roster-missing-skill · Multi-phase orchestration with agent parallelism and structured verification.
untested
Improve toolkit with continuous evaluationworkflowEngineeringL4
toolkit-improvement · Multi-phase orchestration with agent parallelism and structured verification.
untested
Brainstorm with cross-specialist reviewworkflowOpsL2
wf-brainstorm · Synthesizing design decisions across conflicting expert perspectives.
untested
Execute work with deterministic verificationworkflowEngineeringL3
fierce-ralph · Iterative task execution with real-time progress verification and stagnation detection.
untested
Screen EHR products with parallel agentsworkflowProductL2
cheap-pass-100 · High-volume batch screening with fixed concurrent worker capacity.
untested
Debug resume parsing across model variantsworkflowEngineeringL2
trace-resume-model · Distinguishing prefix vs. content-addressed workflow resumption models.
untested
Ship code through architecture review gatesworkflowEngineeringL2
feature-dev · Multi-gate feature development with human approval at requirements and architecture phases.
untested
Refactor neural net primitives with bit-level verificationworkflowEngineeringL3
finish-l1-primitives · Refactoring monolithic code into bit-identical primitives with gated extraction.
untested
Generate and verify files at scale with metricsworkflowOpsL2
parallel-file-write-with-verify · Writing many files in parallel then auditing for quality metrics.
untested
Gate releases with adversarial quality reviewworkflowEngineeringL2
gate · Code review with parallel specialists and adversarial skepticism.
untested
Regenerate specs across voice styles automaticallyworkflowProductL3
spec-critique-voice-ladder-regen · Criticizing specifications from multiple evaluation angles with voice consistency.
untested
Audit code for correctness and deploy fixesworkflowEngineeringL2
go-audit-fix · Automated auditing and fixing of Go code issues.
untested
Map content type coverage and edge casesworkflowProductL2
type-field-coverage · Measuring type coverage across large codebases.
untested
Implement specs autonomously end-to-endworkflowEngineeringL3
sdd-implement-engine · Implementing complex systems driven by formal specifications.
untested
Safely triage and quarantine untrusted inputsworkflowOpsL2
triage_quarantine · Isolating failing cases for targeted root-cause investigation.
untested
Push completed work and open pull requestworkflowEngineeringL2
docking-workflow · Integrating components with automatic verification.
untested
Map codebase architecture in parallelworkflowEngineeringL2
fd-understand-codebase · Building knowledge of unfamiliar codebases through structured exploration.
untested
Investigate issues and synthesize findingsworkflowEngineeringL2
assess-investigate · Code review with parallel specialists and adversarial skepticism.
untested
Port database layer with integration testsworkflowEngineeringL3
phase1-repos · Processing multiple repositories in a phased workflow.
untested
Validate research manuscript for publicationworkflowL2
g3-readiness-review · Rigorously vetting academic manuscripts through multiple independent reviewer lenses before submission, with explicit synthesis of conflicts and prioritized fixes.
untested
Plan agent memory implementation strategyworkflowEngineeringL3
fathomdb-agent-memory-impl-strategy · Grounding implementation strategy for agent-memory gaps in the actual engine code, with adversarial re-verification against real mechanisms and invariants.
untested
Reverse-engineer trading strategies from resultsworkflowDataL3
polymarket-reverse-engineer · Designing independent, first-principles trading strategies from observed edge patterns, falsifiable through rigorous gauntlet testing, without relying on wallet-copying.
untested
Validate and land compiler features systematicallyworkflowEngineeringL3
drain-wave-1 · Draining large spec items with sub-agent isolation and orchestrator-side independent verification of runtime results, preventing sub-agent report blindness.
untested
Assess drug discovery feasibility scientificallyworkflowL2
insilico-feasibility-and-track-record · Honestly evaluating which in-silico predictions have real clinical track records versus hype, with explicit real_world_status per example.
untested
Scan and categorize edge cases systematicallyworkflowEngineeringL2
phase1e-scan · Surfacing external unknowns that could block Phase 2 launch through execute-access probing (RUN code, not just read), with optional skeptic verification gated on FAIL only.
page 131 / 162