e-volv

Agent Ops playbook

PromptRegressionSuite

Weekly re-run of golden tasks against fixed rubrics, tracking quality drift, cost, and latency per agent and playbook.

Weekly, re-runs a fixed set of golden tasks — a known PR against code-review, a known ticket against break-down-ticket, and a known failure against fix-ci — scores each output against a rubric and the previous week's output, flags quality drift and attributes it to prompt edits, model version changes, tool changes, or temperature changes, and tracks cost and latency per task so a quality gain that cost 3× is not mistaken for an improvement.

Identifier
prompt-regression-suite
Version
1.0.1
Steps
4
Triggers
1
  • prompt
  • regression
  • quality
  • agent-ops
  • golden-tasks
  • monitoring

When it runs

Scheduleschedule.weekly
Schedule
Every Monday at 06:00 UTC 0 6 * * 1
Timezone
UTC

The pipeline

The graph below is the one the workflow opens with in the builder — same steps, same layout, drawn on the same canvas. The run playing through it is a simulation; the branches and conditions are real.

  1. 01
    Weekly Schedule (Monday 06:00 UTC)trigger

    The event that starts the run.

  2. 02
    Clone Repository ⚠️ SET YOUR REPOgit.clone

    Shallow-clones the repository at the right ref.

  3. 03
    Run Golden Task Regressionagent.session

    A coordinator delegates to specialist agents.

  4. 04
    File Regression Ticket ⚠️ SET YOUR TICKET INTEGRATIONticket.create

    Opens a ticket on the connected tracker.

The agent

Prompt Regression Coordinator

Base type
Senior Developer
Temperature
0.2
Max iterations
50
Tools
7

Memory · 2

  • memory_writeWrite Memory · write
  • memory_readRead Memory · read

Filesystem · 1

  • read_fileRead File · read

Terminal · 1

  • run_terminal_cmdRun Terminal Command · write

Code search · 1

  • code_searchCode Search · read

Tickets · 1

  • create_ticketCreate Ticket · write

Status · 1

  • update_statusUpdate Status · write