Specification clarity
An agent builds the wrong thing faithfully, fast, and at scale. Ambiguity is now a throughput problem.
The operating system for engineering programs
AI made building fast. Knowing what to build, proving it works, and getting it accepted did not get faster. Progros turns your requirements into one live graph — from what the customer asked for to the evidence that proves you delivered it.
Start with one requirements document. No migration, no process change.
01 Why now
An agent can turn a precise, testable spec into working, tested code in minutes. So the constraint moved — to knowing what to build, proving it works, and getting it accepted. That is the work Progros manages.
An agent builds the wrong thing faithfully, fast, and at scale. Ambiguity is now a throughput problem.
Test benches, device labs, hardware, qualified witnesses, customer experts. None of it got faster.
Output volume rose. The rate at which customers absorb, train on, and accept change did not.
What should be true, what risk is acceptable, what evidence is enough. Not delegable — and never delegated.
02 What it costs
The requirement was ambiguous. The agent didn't ask. Nobody noticed until the demo.
The results are real. Which ones still hold, nobody can say.
The binding constraint was physical, and invisible to every tool in use.
The customer wouldn't operate a capability nobody had agreed how to accept.
Every one of them was knowable in week two — by a tool that modeled intent, evidence, and reality.
03 The structural gap
Done5 ptsSprint 14test-results-final-v2.xlsx
9f1c2ab7 on main · evidence currentYour tracker isn't bad at this. It's structurally incapable of it.
04 The diagnostic
"comply with GDPR"CriticalExplicit consent, a mandatory impact assessment, rights to access and erasure. Three words expand to a quarter of the real scope.
"99.9% uptime"CriticalIt needs weeks of production observation. A launch gate of "all requirements verified" is arithmetically impossible.
(not in the document)CriticalPeople log workouts in basements and on trails. No line-by-line review finds a requirement that isn't there.
"during business hours"HighFitness apps peak mornings, evenings, and weekends — precisely outside the window. Whose time zone, anyway?
A line-by-line review reads the text. Progros checks what the text implies — and what it leaves out. Minutes after upload, with every claim citing the line it came from.
05 How it works
One graph, built by agents that propose and people who decide. Every screen and report is a view of it — so none of them can drift.
Upload a spec, statement of work, or feature list. The Elicitor proposes atomic, testable claims, each quoting the exact line it came from. A claim that can't cite its source is discarded.
The Analyst reads the whole specification: implied obligations, unverifiable claims, conflicts, gaps. You acknowledge, resolve, or dismiss with a reason — and dismissals stay dismissed.
Intent lives across the whole spec, not in any one line. Progros derives the customer's intent, end-to-end acceptance scenarios, and criteria for every requirement — and asks the customer instead of inventing a number.
Coding agents connect over MCP and read each requirement's acceptance contract before they write a line. When the spec is unclear, they ask your team rather than guess.
Tests name the requirement once. Every CI run after that lands as evidence pinned to the exact commit: passing moves a requirement to Verified, a failure moves it back, and a newer build that didn't re-test it marks the proof stale.
One page tells a program lead what is proven, what is failing, what went stale, and which customer scenarios are ready to demo — live, from the evidence, never from a status meeting.
06 For engineers
One annotation, written once, in the test you were writing anyway. Every CI run produces attributed evidence pinned to the commit it proves. Commits and CI advance the lifecycle; people make only the judgment calls.
Conventions are encouraged — never a merge gate. Bookkeeping must never block shipping.
// auth.spec.ts — the test you were writing anyway
test(
'session expires after 30 min idle @verifies("R-1.3")',
async () => {
// …
});
$ node progros.mjs report-tests results/*.xml
Progros: 214 tests (213 passed, 1 failed) at 9f1c2ab7e0d4
✓ R-1.3 2 test(s)
✓ R-2.1 4 test(s)
✗ R-3.4 3 test(s), 1 failed
→ R-3.4 verified → implemented
⚠ unknown requirement refs: R-9.9 — check the annotations
Works with any CI that writes JUnit XML — Bitbucket, GitHub, GitLab, and the rest.
07 Governed AI
Narrow roles, not one chatbot — an Elicitor, an Analyst, a Verifier — each with a registered, versioned prompt.
Every proposal goes through review: accept, amend, or reject. How much agents may do on their own is set by the program's rigor.
Every contribution records who or what made it, which model and prompt, and who approved it — on a tamper-evident, hash-chained log. Your data never trains a model.
› Show everything in this baseline that came from an AI proposal, and who approved it.
| Claim | Origin | Decision | By |
|---|---|---|---|
| R-1.3 | Elicitor | Approved | Security lead |
| R-8.11 | Analyst | Approved with edits | Data protection officer |
| R-9.2 | Verifier | Rejected | Engineering lead |
| R-3.7 | Elicitor | Approved | Product owner |
214 of 312 claims agent-originated · 100% human-reviewed · 0 unattributed
08 The rigor dial
Rigor profiles and evidence-driven lifecycle are available today; e-signature and witnessed verification arrive with the regulated tier.
Run a prototype at Lean and a certified subsystem at Safety-critical — same people, same benches, same tool.
Go from your first customer to your first certification by turning a dial, not by switching to a heavyweight ALM.
Tenant isolation enforced by the database, immutable baselines, and a hash-chained audit log — not bolted on later.
09 Every engineering discipline
Progros doesn't mirror your repository or your PDM vault. It records which revision of which work product proves which requirement — the same model whether that's a commit, a board revision, or a stress report.
10 Where it fits
| Intent & requirements | Evidence & verification | Acceptance from intent | Governed AI | Engineer-grade UX | Resources & plans against reality | |
|---|---|---|---|---|---|---|
| Issue trackersJira · Linear · Azure DevOps | ||||||
| Requirements / ALMDOORS · Polarion · Codebeamer · Jama | ||||||
| Test managementTestRail · Zephyr · qTest | ||||||
| AI coding agentsMake the new bottleneck worse, faster | ||||||
| ProgrosBy design |
Core strengthPartialStructurally absentOn our roadmap
11 Getting started
Upload a requirements document. Within minutes you have cited claims and a diagnostic of what's missing, implied, or unprovable.
Review the derived intent, scenarios, and criteria — then take the open questions to your customer before anyone builds.
Add one step to your pipeline and connect your coding agents. Evidence starts flowing from the next build.
Regulated customer, certification, government contract: turn the dial for that program. Same tool, same data.
Get started
Create a workspace in a couple of minutes and run your first diagnostic today.
Already have an account? Sign in. Questions first? Read the FAQ.
Received
We'll reply to about your program. Meanwhile you can create a workspace and try a spec.
Questions
Not on day one, and you don't need it to. Start with a diagnostic of one specification and keep your tracker. Progros holds what a tracker can't — what must be true, how it's proven, and whether the customer accepts it — and work can be derived from the gap between intent and evidence.
No. Your data is never used to train any model. Every model call is recorded on a ledger with its inputs, outputs, model, and prompt version, so you can see exactly what the AI did with what.
Claude, by Anthropic, through Progros's own inference gateway — which records every call, attributes it to the program it served, and keeps the agents' prompts registered and versioned.
Agents propose; people decide. Every claim, criterion, and scenario an agent drafts goes to a review queue where you accept, amend, or reject it. At higher rigor levels, agents may only propose. Coding agents connected over MCP can read requirements and ask questions — they can never edit the specification.
It was designed for you from the first line: database-enforced tenant isolation, immutable baselines, a hash-chained audit log, and rigor profiles up to Defense. The architecture is built so dedicated, government-cloud, and air-gapped deployments run the same code; e-signature and witnessed verification arrive with the regulated tier. Tell us about your program when you sign up and we'll talk through what you need.
Software is connected today, and firmware and lab tests that produce JUnit results work through the same path. Electrical, mechanical, structures, thermal, systems, and manufacturing use the same model — a source, its revisions, and the evidence they produce — and connectors for them are on the roadmap. Tell us what you use.