README: state the regulated-workflow capability and OIC development hypothesis - #34
Conversation
inventor1975
left a comment
There was a problem hiding this comment.
Approving, bound to exact head 79d8edd90934638d58e06e7ad5700cbe93c507a3.
Composition confirmed against the work order before reading: PR open, head and base exactly as specified, one changed file (README.md), CI run 32904839488 SUCCESS. No branch edit, push, rebase, merge or wording change made during this review.
All twelve review questions verified against the README text at that head:
Claim classification (Q4, Q6, Q7, Q8, Q10) — clean.
All eight full-benchmark thresholds carry TARGET - NOT MEASURED, individually, not as a blanket caveat. The human-efficiency threshold is separated and carries the stronger PROVISIONAL TARGET - NOT MEASURED - NOT CALIBRATED. Safe Automation Coverage reads PROPOSED - NOT YET PART OF THE FROZEN OIC-BENCH SPECIFICATION, and is described in the conditional ("would measure"), not the indicative. Semantic OIC-Bench preflight and Full OIC-Bench v0.1 both read NOT YET RUN. Semantic implementation remains stated as blocked.
Evidence separation (Q5, Q9) — clean and explicit. The README defines MEASURED as an observed result tied to a reproducible run, and states plainly that the current results verify repository infrastructure and are not semantic OIC-Bench evidence. That sentence is the load-bearing one, and it is present rather than implied.
Boundary (Q2, Q3) — preserved. "Runtime authorization starts downstream" is stated as a heading, not buried. OPA is described as evaluating policy and returning decisions to the integrating application, with Rego named as OIC's first proposed executable target. Nothing in the text positions OIC as a replacement for OPA, Cedar or runtime authorization; comparable systems are acknowledged as mature.
Roadmap (Q11) — evidence-gated as required. "No stage is complete because its features exist. Advancement requires the named benchmark, verification, operational, and independent-review evidence." Stage 0 is CURRENT with a stated exit condition; the gate is not schedule-driven.
No status promotion (Q12). No license grant is asserted — licensing is explicitly pending. No maturity, production-readiness or runtime-validity claim appears.
One observation, not a requested change: the discipline in this README is stronger than the norm for repository front pages, specifically because the eight thresholds are each marked rather than covered by a single disclaimer at the top. A blanket caveat is easy to read past; per-row marking is not. That choice is worth keeping as the file grows.
Boundary of this review, stated plainly: I verified the text at the exact head and the composition metadata via the API. I did not execute the repository's tests, did not evaluate whether the stated targets are the right targets, and did not assess whether the underlying capability claims are achievable. This review is about whether the document distinguishes claimed from measured. It does.
— Reviewed by Claude (Opus 5) on Vitaly Reznik's standing review authorisation. Merge remains Vitaly Reznik's decision, and he is not awake; nothing here authorises a merge.
Documentation objective
Present the regulated-enterprise capability thesis, falsifiable benchmark
targets, explicit evidence states, named comparative baselines, proposed future
automation-coverage metrics, and an evidence-gated development roadmap.
This final correction separates proposed full-benchmark gates from the
provisional human-efficiency comparative-claim threshold and uses this PR as the
candidate composition evidence record.
Exact candidate composition
d06917fa6877277d7118b49e80d6a69446f50712e307425144eeb6dc9dc4d613b35829222630891079d8edd90934638d58e06e7ad5700cbe93c507a35d34078434678b4dbf89b0f497c2fad77794d89e32904839488README.mdonlyEvidence classification
Proposed full-benchmark targets remain separately classified and not measured:
source support, unsupported-field rate, unknown-to-false, authority
reconstruction, ambiguity recall, false-resolution, behavioral conformance,
and change-impact recall.
The human-efficiency comparison is separately classified as:
PROVISIONAL TARGET - NOT MEASURED - NOT CALIBRATED. It is not an absolutesafety or release gate.
Safe Automation Coverage and admitted automation yield under fixed
expert-review budget remain proposed future metrics, not formally adopted into
OIC-Bench.
Composition-bound infrastructure measurements
The following measured counts bind to final branch head
79d8edd90934638d58e06e7ad5700cbe93c507a3, based06917fa6877277d7118b49e80d6a69446f50712, and GitHub PR CI run32904839488, which executed synthetic merge composition5d34078434678b4dbf89b0f497c2fad77794d89e:INCOMPLETE, required exit 3.BLOCKED.NOT YET RUN.NOT YET RUN.These measurements verify repository infrastructure. They are not semantic
OIC-Bench results and are not evergreen project statistics.
GitHub Actions results
All nine applicable jobs in run 32904839488 concluded success:
GitHub Actions evaluated the synthetic merge composition identified above; this
is not characterized as raw exact-head execution.
Boundaries preserved
runtime, network, or telemetry change.
Required declarations