The operating system for engineering programs

The system of record for
verified intent.

AI made building fast. Knowing what to build, proving it works, and getting it accepted did not get faster. Progros turns your requirements into one live graph — from what the customer asked for to the evidence that proves you delivered it.

Start with one requirements document. No migration, no process change.

  1. Need
  2. Requirement
  3. Capability
  4. Work
  5. Evidence
  6. Training
  7. Acceptance

01 Why now

AI made building fast. Nothing around it got faster.

An agent can turn a precise, testable spec into working, tested code in minutes. So the constraint moved — to knowing what to build, proving it works, and getting it accepted. That is the work Progros manages.

01

Specification clarity

An agent builds the wrong thing faithfully, fast, and at scale. Ambiguity is now a throughput problem.

02

Verification capacity

Test benches, device labs, hardware, qualified witnesses, customer experts. None of it got faster.

03

Stakeholder trust

Output volume rose. The rate at which customers absorb, train on, and accept change did not.

04

Judgment

What should be true, what risk is acceptable, what evidence is enough. Not delegable — and never delegated.

02 What it costs

Four failures every program lead has lived.

  1. 01

    We built the wrong thing — fast.

    The requirement was ambiguous. The agent didn't ask. Nobody noticed until the demo.

  2. 02

    "Verified" — three architecture changes ago.

    The results are real. Which ones still hold, nobody can say.

  3. 03

    Month five: the test bench is booked until month nine.

    The binding constraint was physical, and invisible to every tool in use.

  4. 04

    Shipped. Refused.

    The customer wouldn't operate a capability nobody had agreed how to accept.

Every one of them was knowable in week two — by a tool that modeled intent, evidence, and reality.

03 The structural gap

A ticket records activity.
A claim records what's true.

PRJ-1423 · StoryIssue tracker

Implement session timeout

Done5 ptsSprint 14test-results-final-v2.xlsx

  • Which requirement does this satisfy?
  • Proven on which build — and is that still true?
  • Which customer need does it serve?
  • Has the customer agreed what "done" means?
R-1.3 · Requirement · Baseline v4Progros

Sessions expire after 30 minutes of inactivity.

Serves
N-2 · Users distrust apps holding health data
Accepted when
Given an idle session, after 30 minutes the next request is refused · agreed with the customer
Proven by
2 tests at 9f1c2ab7 on main · evidence current
Method
Automated test · @verifies in auth.spec.ts
Scenario
C-2 Returning user on a shared device
Provenance
Proposed by the Elicitor agent · approved by the security lead

Your tracker isn't bad at this. It's structurally incapable of it.

04 The diagnostic

Each one looked like a finished requirement.

Input · a fitness app, verbatim · 12 lines
  1. User Authentication – Secure login with email/password and optional social login.
  2. Goal Setting – Users can create and update fitness goals.
  3. Activity Tracking – Log workouts, distance, and calories burned.
  4. Progress Reports – Generate weekly/monthly summaries.
  5. Notifications – Send reminders for workouts and milestones.
  6. Data Export – Export activity data in CSV format.
  7. Performance: App must load within 2 seconds on mobile devices.
  8. Security: Data encrypted at rest and in transit; comply with GDPR.
  9. Reliability: 99.9% uptime during business hours.
  10. Usability: Intuitive UI; accessible to users with basic literacy.
  11. Scalability: Support up to 10,000 concurrent users.
  12. Must be developed using React Native.
"comply with GDPR"Critical

Health data is special-category data.

Explicit consent, a mandatory impact assessment, rights to access and erasure. Three words expand to a quarter of the real scope.

"99.9% uptime"Critical

Unverifiable before launch. Ever.

It needs weeks of production observation. A launch gate of "all requirements verified" is arithmetically impossible.

(not in the document)Critical

Offline logging is missing entirely.

People log workouts in basements and on trails. No line-by-line review finds a requirement that isn't there.

"during business hours"High

The SLO is silent when it matters.

Fitness apps peak mornings, evenings, and weekends — precisely outside the window. Whose time zone, anyway?

A line-by-line review reads the text. Progros checks what the text implies — and what it leaves out. Minutes after upload, with every claim citing the line it came from.

05 How it works

From the customer's words to proof they'll accept.

One graph, built by agents that propose and people who decide. Every screen and report is a view of it — so none of them can drift.

  1. 01

    Extract

    Upload a spec, statement of work, or feature list. The Elicitor proposes atomic, testable claims, each quoting the exact line it came from. A claim that can't cite its source is discarded.

    R-7 The system shall let users log workouts, distance, and calories.“Log workouts, distance, and calories burned.” · line 3
  2. 02

    Diagnose

    The Analyst reads the whole specification: implied obligations, unverifiable claims, conflicts, gaps. You acknowledge, resolve, or dismiss with a reason — and dismissals stay dismissed.

    Critical Implied obligation · GDPR Art. 9 health data targets R-12, CON-1 · 2 proposed requirements
  3. 03

    Agree what "done" means

    Intent lives across the whole spec, not in any one line. Progros derives the customer's intent, end-to-end acceptance scenarios, and criteria for every requirement — and asks the customer instead of inventing a number.

    Given no connectivity, a logged workout is kept and syncs within [TBD: maximum sync delay]Question for the customer · R-26
  4. 04

    Build — with your agents

    Coding agents connect over MCP and read each requirement's acceptance contract before they write a line. When the spec is unclear, they ask your team rather than guess.

    get_requirement R-7 → statement · criteria · traces to N-1 · scenario C-9 · open questions
  5. 05

    Prove it

    Tests name the requirement once. Every CI run after that lands as evidence pinned to the exact commit: passing moves a requirement to Verified, a failure moves it back, and a newer build that didn't re-test it marks the proof stale.

    ✓ R-1 2 tests · ✗ R-3 1 failed · ⚠ R-999 unknown ref
  6. 06

    Know where you stand

    One page tells a program lead what is proven, what is failing, what went stale, and which customer scenarios are ready to demo — live, from the evidence, never from a status meeting.

06 For engineers

Engineers never open it to update a status.

One annotation, written once, in the test you were writing anyway. Every CI run produces attributed evidence pinned to the commit it proves. Commits and CI advance the lifecycle; people make only the judgment calls.

  • Proposed agent drafts
  • Baselined you decide
  • Allocated inferred
  • Designed inferred
  • Implemented inferred
  • Verified from evidence
  • Validated you decide
  • Accepted the customer signs

Conventions are encouraged — never a merge gate. Bookkeeping must never block shipping.

// auth.spec.ts — the test you were writing anyway
test(
  'session expires after 30 min idle @verifies("R-1.3")',
  async () => {
    // …
  });
$ node progros.mjs report-tests results/*.xml
Progros: 214 tests (213 passed, 1 failed) at 9f1c2ab7e0d4
  ✓ R-1.3  2 test(s)
  ✓ R-2.1  4 test(s)
  ✗ R-3.4  3 test(s), 1 failed
  → R-3.4 verified → implemented
  ⚠ unknown requirement refs: R-9.9 — check the annotations

Works with any CI that writes JUnit XML — Bitbucket, GitHub, GitLab, and the rest.

07 Governed AI

AI you can hand to an auditor.

  1. 01

    Agents propose.

    Narrow roles, not one chatbot — an Elicitor, an Analyst, a Verifier — each with a registered, versioned prompt.

  2. 02

    People baseline.

    Every proposal goes through review: accept, amend, or reject. How much agents may do on their own is set by the program's rigor.

  3. 03

    Provenance is permanent.

    Every contribution records who or what made it, which model and prompt, and who approved it — on a tamper-evident, hash-chained log. Your data never trains a model.

Program Atlas · Baseline v4Illustrative

› Show everything in this baseline that came from an AI proposal, and who approved it.

ClaimOriginDecisionBy
R-1.3ElicitorApprovedSecurity lead
R-8.11AnalystApproved with editsData protection officer
R-9.2VerifierRejectedEngineering lead
R-3.7ElicitorApprovedProduct owner

214 of 312 claims agent-originated · 100% human-reviewed · 0 unattributed

08 The rigor dial

One product, from prototype to defense program.

LeanPrototype CommercialMost teams start here RegulatedMedical · fintech Safety-criticalAuto · aerospace Defense programGovernment
Change controlAuthor approves Quorum board, e-signature
VerificationImplementer may verify Independent and witnessed
FindingsAdvisory Gate-blocking
Lifecycle from evidenceCI moves claims to Verified Evidence recorded, a person attests

Rigor profiles and evidence-driven lifecycle are available today; e-signature and witnessed verification arrive with the regulated tier.

Per program

Run a prototype at Lean and a certified subsystem at Safety-critical — same people, same benches, same tool.

Raise without migrating

Go from your first customer to your first certification by turning a dial, not by switching to a heavyweight ALM.

Built for it from day one

Tenant isolation enforced by the database, immutable baselines, and a hash-chained audit log — not bolted on later.

09 Every engineering discipline

Link versions, not copies.

Progros doesn't mirror your repository or your PDM vault. It records which revision of which work product proves which requirement — the same model whether that's a commit, a board revision, or a stress report.

  • SoftwareAvailable
  • Firmware & embeddedTests via JUnit today
  • Test & labTests via JUnit today
  • ElectricalPlanned
  • MechanicalPlanned
  • StructuresPlanned
  • Thermal & fluidsPlanned
  • Systems (MBSE)Planned
  • Manufacturing & qualityPlanned

10 Where it fits

Everything exists somewhere. Nothing in one graph.

How Progros compares with the tool categories programs use today
Intent & requirementsEvidence & verificationAcceptance from intentGoverned AIEngineer-grade UXResources & plans against reality
Issue trackersJira · Linear · Azure DevOps
Requirements / ALMDOORS · Polarion · Codebeamer · Jama
Test managementTestRail · Zephyr · qTest
AI coding agentsMake the new bottleneck worse, faster
ProgrosBy design

Core strengthPartialStructurally absentOn our roadmap

11 Getting started

Start with a diagnostic. Change nothing.

  1. 1

    Bring one spec

    Upload a requirements document. Within minutes you have cited claims and a diagnostic of what's missing, implied, or unprovable.

  2. 2

    Agree acceptance

    Review the derived intent, scenarios, and criteria — then take the open questions to your customer before anyone builds.

  3. 3

    Connect the work

    Add one step to your pipeline and connect your coding agents. Evidence starts flowing from the next build.

  4. 4

    Raise rigor when you need it

    Regulated customer, certification, government contract: turn the dial for that program. Same tool, same data.

Get started

Bring one spec.
See what it finds.

Create a workspace in a couple of minutes and run your first diagnostic today.

  1. Create your workspace with your work email.
  2. Confirm your email — the link opens Progros, ready to go.
  3. Upload one spec and read what it finds.

Create your workspace

Already have an account? Sign in. Questions first? Read the FAQ.

Regulated, safety-critical, or defense?

Tell us about your program — we'll talk through deployment, rigor, and how your data is handled.

Engineering disciplines involved optional

We use your details only to contact you about Progros. No newsletters, no sharing.

Questions

Asked, answered.

Does Progros replace Jira?

Not on day one, and you don't need it to. Start with a diagnostic of one specification and keep your tracker. Progros holds what a tracker can't — what must be true, how it's proven, and whether the customer accepts it — and work can be derived from the gap between intent and evidence.

Does my data train AI models?

No. Your data is never used to train any model. Every model call is recorded on a ledger with its inputs, outputs, model, and prompt version, so you can see exactly what the AI did with what.

Which AI does it use?

Claude, by Anthropic, through Progros's own inference gateway — which records every call, attributes it to the program it served, and keeps the agents' prompts registered and versioned.

Can the AI change our requirements?

Agents propose; people decide. Every claim, criterion, and scenario an agent drafts goes to a review queue where you accept, amend, or reject it. At higher rigor levels, agents may only propose. Coding agents connected over MCP can read requirements and ask questions — they can never edit the specification.

We're regulated, or a defense program. Is it ready for us?

It was designed for you from the first line: database-enforced tenant isolation, immutable baselines, a hash-chained audit log, and rigor profiles up to Defense. The architecture is built so dedicated, government-cloud, and air-gapped deployments run the same code; e-signature and witnessed verification arrive with the regulated tier. Tell us about your program when you sign up and we'll talk through what you need.

Is it only for software?

Software is connected today, and firmware and lab tests that produce JUnit results work through the same path. Electrical, mechanical, structures, thermal, systems, and manufacturing use the same model — a source, its revisions, and the evidence they produce — and connectors for them are on the roadmap. Tell us what you use.