https://www.millabs.net/blog/openproject-evidence-pipeline-fedramp-20x/

This is a follow-on to Deploying OpenProject Into a Hardened DoW Enclave. Everything here was produced in the same private Millabs lab, which is not a government-authorized environment. It is a demonstration, not a FedRAMP package.

In August I reported that OpenProject's two container images went from 685 known vulnerabilities to zero after hardening. The platform's image gate scanned both builds on August 17 and signed the result into an attestation on each image: zero critical, zero high, zero medium, zero low.

FedRAMP's rules for 2026 ask a different question. Not "was it clean when you scanned it?" but "is it clean now, and can you show the evidence every day?" So I built the pipeline that answers that question and ran it against our own study.

What changed in FedRAMP

FedRAMP's Consolidated Rules for 2026 (CR26) took effect on July 4, 2026, and every new 20x application must already follow them. From January 1, 2027 they bind every 20x-certified offering and every new Rev5 application, and from June 11, 2027 FedRAMP accepts no new Rev5 applications at all. For the 20x path, three things replace the paperwork most vendors know:

  • The Security Decision Record replaces the System Security Plan. It is a JSON document with a published schema, and it must also be supplied in human-readable form.
  • Key Security Indicators replace control-by-control narratives. Class C (the old Moderate) has 46 of them, in ten themes, and 44 cite related NIST SP 800-53 controls. They ask whether security outcomes are actually happening: is configuration checked for drift, are changes redeployed from version control, are third-party components monitored for new vulnerabilities.
  • Persistent validation replaces the snapshot. Class C requires up to a year of daily metric history for each indicator.

If your buyer is the Department of War, none of this replaces an ATO. DoW systems are still authorized under RMF, by a DoW authorizing official. But the evidence underneath is the same: scans, attestations, manifests, restore tests, reviews. Collect it once and it can feed either.

What we built

The pipeline has three layers.

  1. Collectors read evidence and write it to an append-only log: vulnerability scans, the image gate's signed scan attestations, Kubernetes manifests, static code analysis, cluster state (policy reports, GitOps status, mesh encryption, backups), and human-recorded reviews.
  2. Measures are declared for each indicator: an objective, a threshold and a cycle. Each is scored against its latest evidence as pass, fail, stale (it passed, but the evidence is older than its cycle allows) or no evidence.
  3. The record is generated from decision records a person writes (who owns each indicator, its status, and the gap and its risk if it is not met) and validated against FedRAMP's published schema on every build. The JSON and the human-readable view are generated from the same model, so they cannot disagree. A build that fails validation writes nothing.

Every piece of evidence says how it was obtained: observed live from a running system, observed static from an artifact in version control, or documented by a person from a report. That label turned out to matter more than anything else in the record.

What the evidence said

The zero lasted 52 days. The pipeline read the gate's signed attestations, verified them against the platform's key, and then rescanned the exact digests staging was running.

Image August 17, signed attestation October 8, same digest
Rails 0 28 (7 high)
Collaborative editor (Hocuspocus) 0 24 (1 critical, 9 high)

Nothing in the images changed. New advisories were published against packages already inside them, mostly in the base image: OpenSSL, BusyBox, Ruby, npm. A scan result is evidence about the day it was scanned. The gate scans once, at promotion, and its signed attestation cannot learn anything afterward. That is by design, since an attestation you could rewrite would be worthless, but it means someone has to rescan what is running.

The fix is a rebuild, and the first rebuild failed. The hardened Ruby base had moved to Ruby 4.0.7, itself a security release. OpenProject pins Ruby 4.0.6 exactly, so the build refused to run. An image you cannot rebuild is an image you cannot patch. Moving the pin was a one-line change.

Getting that one line merged surfaced a second problem. The registry credentials are protected so that only the main branch can push images, which is right. Every branch pipeline therefore failed at login, so no change could pass a reviewed merge request, which is also a reasonable rule. Each control was sound; together they meant a fix could not reach the main branch through review. Branch pipelines now build without pushing, which is the check a review needs, and only the main branch pushes.

After the rebuild, staging runs the new images, verified at the cluster and over HTTPS:

Image Rebuilt October 8 What remains
Rails 8 (2 high) All in two gems the application pins: dalli (fixed in 5.0.10) and rubyzip (fixed in 3.4)
Collaborative editor 3 (2 high) All in two editor libraries the application pins: @tiptap/core and prosemirror-view

Every base-image finding cleared. The remainder sat in application dependencies, which no base image reaches. Each one needs either a dependency update or a written security impact statement that a platform's security owner will accept. Updating is better when it can be done safely, so we did that.

Mitigating what remained

Where upstream OpenProject had already made a fix, we ported it rather than inventing our own. Where it had not, we pinned fixed versions and held ourselves to the project's own supply-chain rules: OpenProject holds new releases back until they have aged: its Gemfile tells Bundler to refuse any gem version younger than seven days, and its Dependabot configuration waits 30, 14 and 5 days before proposing major, minor and patch npm updates.

Finding Where Mitigation
rubyzip path traversal (high) Rails Version 3.4.1, ported from upstream's 17.9 release. Our older source also needed one line changed: version 3 altered the call the backup job uses. Upstream had rewritten that file in September; we found the break by reading and proved the fix by running the call.
dalli command injection (high) Rails Version 5.0.7.
dalli, six further advisories (unrated) Rails Waiting. The fixed releases are two to six days old, inside OpenProject's own seven-day cooldown. Two sound controls disagree here: the cooldown protects against a compromised fresh release, and the advisories are real. We kept the cooldown, because the affected client is a memcached client this deployment never connects, and scheduled the update for when the release clears it on October 13.
@tiptap/core ReDoS (high) and attribute pollution (medium); prosemirror-view XSS on paste (high) Collaborative editor server and the browser bundle Every @tiptap package pinned to 3.31.3 and prosemirror-view to 1.42.5. The editor packages pin each other exactly, so they move together. The browser copy needed one more pin: the newer view expects a newer document model, npm installed a second copy, and the TypeScript build refused two incompatible versions of the same types.

The scanner's blind spot. The image scan found the editor libraries in the collaborative-editing server, where an XSS-on-paste flaw is hardly reachable: a server has no browser and no clipboard. It said nothing about the copy that matters, the one compiled into the JavaScript every user's browser runs. Compiled bundles are invisible to image scanners. So we scanned the lockfile the bundle is built from: 35 known vulnerabilities, one critical and 15 high, and no image scan had reported a single one of them against the bundle. We then checked each against the production build's own list of bundled third-party packages:

Findings Ships to browsers?
Editor libraries (the same three, in the browser's copy) 3 Yes, fixed above
Angular framework (@angular/core sanitization bypass, @angular/common transfer-cache leak) and moment 3 (medium) Yes. Fixed by porting upstream's Angular 22.1 update and moment 2.31.0
Build tooling: the Angular CLI's dev server, its MCP support, YAML and query-string parsers, and similar 29, including the critical and 13 high (28 after the Angular update) No. They run on the build machine and never reach the image or the browser

The bundle now carries no known vulnerabilities. The 28 build-time findings still matter, to whoever runs the build, and they belong on the vendor's toolchain update schedule. They are not a risk to the platform or its users, and a security owner should hear exactly that distinction rather than a single combined count.

After mitigation, staging runs the updated images: the collaborative editor scans clean, the Rails image carries only the six dalli advisories waiting out the cooldown, and the browser bundle is clean. We verified the builds, the deployment, the health checks, the backup job's archive call, the collaborative-editing server's own test suite, and that the Angular application boots in a headless browser. We did not test collaborative editing by two people at once, so that is still a manual check.

What else the evidence showed

I read our own manifests first, and they were wrong. Before checking the cluster, I looked at the reference deployment manifests we ship with the study and concluded that the lab was running the unhardened images. It was not. The manifests were a copy taken before the platform repinned staging to the hardened builds. The cluster was right and the copy was stale. The pipeline makes the same distinction on its own: the platform documents strict mutual TLS between workloads, so I declared that indicator Implemented, and the build flagged it at once as declared but never measured.

Hosting on a platform does not mean inheriting it. The lab provides single sign-on through Keycloak. OpenProject's single sign-on integration requires its Enterprise license. Without it, users sign in with application passwords and the platform's identity service does nothing for them. The indicator for passwordless or phishing-resistant authentication is Not Implemented, and it is the application's gap, not the platform's.

The record

Count
Key Security Indicators 46
Measures defined 23
Pass 10
Fail 5
Stale 0
No evidence yet 8

The five failures are each known and explained: helper containers that set no resource limits in their manifests (the lab supplies defaults at admission), single-replica workloads on a single host, the vendor's own static-analysis findings (one measure each for Semgrep and Brakeman), and the six dalli advisories waiting out the cooldown. Of the eight without evidence, four are cluster measures whose collectors are written but not yet scheduled, and four are organizational reviews nobody has recorded.

Who owns the 46

The decision records assign each indicator an owner. In our lab:

Owner Indicators
Platform 18
Shared between platform and application 10
Application 4
The organization running the service 14

A capable hosting platform can carry a real share of the load. But 28 of the 46 involved the application or the organization running it, and 14 are about the organization alone: training, incident after-action reviews, log review, change-procedure review, recovery objectives, removing a customer's data on request. No platform inherits those. For a five-person company they are not hard, but they need a dated record on a schedule, and most small vendors have never kept one.

FedRAMP's schema has no "inherited" status, only Implemented, Partially Implemented and Not Implemented. If you build on someone else's platform, you will have to say in words which part is theirs.

What to do now if FedRAMP is on your roadmap

  • Rescan what is running, on a schedule. Not the image you built last, and not what your registry says at push time.
  • Rebuild on a schedule, and let your runtime pin move with your base. Otherwise the first security release of your language runtime breaks your build at the moment you need it.
  • Make sure a reviewed fix can reach your main branch. A pipeline that cannot pass without production credentials blocks every merge request.
  • Label your evidence. Know which of your claims are measured, which come from an artifact, and which are someone's word. An assessor will ask.
  • Scan what ships to the browser, not only the image. Scan the frontend lockfile and keep only the packages your build actually bundles.
  • Keep your vendor's supply-chain cooldowns, and schedule the follow-up. A fix that waits a few days on purpose is a decision; one nobody tracks is a gap.
  • Check the running system, not a copy of its configuration.
  • Start the organizational habits now. A quarterly training review and a dated log review cost little. A year of daily history cannot be produced the month before you need it.

Robert Burckner is the founder of Millabs Corporation. This work was performed in a privately operated Big Bang lab that is not a government-authorized environment; it carries no USG authorization and holds no ATO. No government data was used. The Security Decision Record described here is a demonstration, not a FedRAMP certification package, and Millabs is not a FedRAMP Recognized independent assessment service. OpenProject is open-source software licensed under GPL-3.0.

If your product is headed for FedRAMP or a DoW enclave and you want to know where your evidence stands before an assessor does, contact Millabs.

Share this post
Email LinkedIn
GET IN TOUCH WITH US

Find out what the gap between your product and an authorized environment actually looks like.