7 From one-file skill to maintainable rounding workflow
A useful skill can begin as one visible SKILL.md. Production readiness is a later operating state, not a larger folder. Add references, scripts, evaluation fixtures, templates, and release controls only when each one addresses a named weakness and has an owner.
During clinical development, the R scripts used for analysis and reporting may change over time. The developer must repeat the rounding review from Chapter 6. The saved response preserves candidate locations, example outputs, explanations, and a proposed solution, but not a repeatable method: which functions must be searched, which ties must be tested, how coverage is recorded, or when to stop before posting an issue. The draft survived the handoff. The ability to reproduce and evaluate it did not.
Chapter 6 established three business rules, defined the scope of BR-001, and showed how to read an exploratory response as a draft rather than an answer key. It also named four things a prompt cannot provide: coverage, repeatability, separation of fixed checks from judgment, and a named human gate. This chapter turns those gaps into requirements and specifies each stage of the AI-first SDLC from Chapter 4.
The AI-Native SDLC Playbook lesson on skills as institutional knowledge describes a skill as a way to make owned, written knowledge operational. This chapter adapts that idea to Biometrics: BR-001 is the source of truth, the rule owner controls it, and a skill tells an agent how to apply it during a bounded review. The chapter follows two implementations of that idea: the book’s one-file teaching skill and the more comprehensive RConsortium/pharma-skills/rounding example. The first makes the contract easy to see. The second shows how the same contract can be separated into maintained instructions, references, scripts, report assets, and evaluation evidence. Both remain advisory. A person still approves the rule and classifies each finding.
The book distributes an instruction-only skill and a small guided exercise. The skill has not been evaluated across repeated agent runs, multiple products, or representative R packages. The results quoted at the pinned metalite.ae commit were verified separately and remain answer-key material, not output from the skill. This is a design pattern, not a reproducible or production-qualified workflow.
The R Consortium repository now contains a broader three-rule implementation with scripts, references, report assets, and benchmark fixtures. It is a useful case study in maintainable packaging, but its repository disclaimer describes the outputs as drafting aids that require independent verification (source). Its presence does not raise this chapter’s maturity level or establish production or GxP readiness for a particular organization.
The chapter has one main finding: start with one inspectable instruction, then promote it through evidence, ownership, and operating controls rather than through file count alone.
- Problem: A useful one-off review cannot be repeated reliably when its procedure remains in one person’s memory.
- Learning objective: Understand how to build a one-file skill that preserves the rule, minimum procedure, coverage, evidence, and accountable human decision, then evolve it into a maintainable skill package when evaluation and operating needs justify the added structure.
- Workflow boundary: Begin with named R source or a pasted R example, BR-001, and the required precision; end when a named human classifies each draft finding.
- Non-goals: Build an agent platform, validate statistical methods, replace independent QC, automate human classification, or claim that either example is qualified for production or GxP use. The one-file lab does not cover BR-002 or BR-003, schedule reviews, or publish issues.
- Why this is well defined: One rule, one bounded source set, observable results, explicit boundaries, versioned inputs when available, and a named decision maker keep the task bounded.
7.1 The evolution strategy
The two examples are not competing designs. They are different stages in one development path.
| Stage | Primary question | Typical structure | Evidence needed to advance |
|---|---|---|---|
| One-file prototype | Can an agent follow the bounded task contract and produce a useful draft? | SKILL.md |
A small answer key, observed runs, and human review |
| Maintainable skill package | Can several contributors change rules, deterministic checks, and report contracts without losing traceability? | SKILL.md plus focused references/, scripts/, assets/, and evals/ |
Versioned rules, executable checks, fixtures, an independent answer key, and regression results |
| Production usage | Can an organization operate the skill safely and consistently for its intended use? | The maintained package plus approved release, access, monitoring, incident, and retirement controls | Accepted benchmarks in the target environment, named ownership, controlled distribution, operating evidence, and rollback criteria |
More files do not move a skill to the next stage. A single file can remain the right design for a small, low-risk advisory task. A comprehensive package can still be unready when its benchmark has not been accepted, its owner is unclear, or its operating controls are missing. The team should stop at the least complex structure that satisfies the approved task contract and evidence requirements.
7.2 Plan — Frame the workflow
7.2.1 Turn the prompt’s limits into requirements
The four gaps in Chapter 6 are not additional metalite.ae defects. They are limitations of the one-off review, and each one becomes a workflow requirement:
- replace reviewer memory with a written minimum procedure while leaving room for the agent to investigate beyond it;
- report coverage for every run, including rules left unchecked;
- keep fixed, repeatable checks separate from exploratory agent work and judgment in the output; and
- name the human who classifies each finding, and stop before any external action.
The function catalog is a required starting point, not a ceiling. The agent can follow wrappers, helper functions, or other unexpected code paths, but it must label candidates found through that exploratory pass separately.
Planning must also name the intended operating level. A one-person experiment, a shared team package, and a routine production control need different evidence, ownership, distribution, and support. Starting with one file is useful only when the team also knows which observed failures would trigger a more maintainable design.
7.2.2 Canonical artifact: Workflow brief
- Objective: Specify a check that finds calls able to round a displayed value or a value used in a report decision, tests them against the half-away-from-zero rule, and can be performed by a reviewer who did not design it.
- Inputs: Named R source or a pasted snippet, its version when available, BR-001 v1.0, the precision used by each relevant call, and the call’s intended report use.
- Intended output: A versioned draft review report containing candidate locations, triage reasons, probe results, coverage, unchecked rules, and items requiring a human decision.
- Baseline evidence: Issue #249 shows what an open investigation can contribute. The manually reviewed source inventory and tie probe below provide the answer key, not measured workflow performance.
- Value hypothesis: A stable minimum check plus an exploratory pass will reduce repeated manual search while keeping the first instruction easy to inspect. Modular resources added in response to measured gaps will make later changes easier to own and test.
- Agent contribution: Search the required catalog and unexpected paths, read surrounding code, run isolated probes, explain evidence, and draft the report. Label observations and judgment separately.
- Human decision: The rule owner approves the procedure and classifies each result as a defect, documented convention, accepted exception, or not applicable.
- Success measures: Select the skill for applicable requests and not for unrelated requests; find every deliberately noncompliant test call; flag no compliant test example; attach rule-specific probe evidence to every finding; state what was not checked; and distinguish candidates, observed results, draft findings, and human decisions.
- Promotion decision: Record whether observed gaps require a reference, deterministic script, report asset, regression fixture, or an operating control, and identify the owner and acceptance evidence for that addition.
- Main risks: A narrow search can produce a convincing false clean report, and a polished skill description can create confidence before its behavior has been evaluated.
The Plan gate is the rule owner’s decision that a repeatable check is worth building. The accepted workflow brief initiates Design.
7.3 Design — Specify the workflow
7.3.1 Carry the rule and its scope clause forward unchanged
BR-001, BR-002, and BR-003, stated in Chapter 6, are separate business rules. BR-001 requires exact decimal ties to round half away from zero and prohibits a displayed negative zero. BR-002 requires calculations to use unrounded values and to round once for display. BR-003 requires display precision to be specified and preserved. Each rule is version 1.0 in this example.
The scope clause attached to BR-001 defines which calls are in scope: any call that can change a numeric value displayed in a report or used to decide which rows the report includes. It therefore places formatC(), sprintf(), and format() among the candidates.
The book prototype checks only BR-001. BR-002 and BR-003 remain explicit gaps, declared in every run rather than left unmentioned. The comprehensive community example expands the contract to all three rules. The reference behavior for BR-001 is stated directly: at zero decimal places, for example, -2.5 and 2.5 become -3 and 3, and displayed zero has no negative sign.
7.3.2 Start with one instruction file
The official Codex skill documentation states that a skill can be one directory containing SKILL.md. Its frontmatter contains name and description; scripts, references, and assets are optional.
The book lab deliberately chooses the smallest visible structure:
rounding/
SKILL.md
The task contract defines what the one file must make visible. The initial lab does not need supporting resources. The book keeps the artifact at rounding/SKILL.md so readers can find it without navigating a hidden installation directory.
The official Codex skill documentation also describes the larger package shape and progressive disclosure: Codex sees the name and description first, loads SKILL.md when the skill applies, and uses optional resources when the instructions direct it to them. This makes SKILL.md the stable entry point even after the package grows. It should route the work, not duplicate every policy, test case, and output template.
7.3.3 Design the maintainable target
At commit 544aa0818208a973e6e22901118c8f97b758c146, the R Consortium example separates the broader workflow this way:
rounding/
SKILL.md
references/
br-001-tie-method.md
br-002-rounding-stage.md
br-003-display-precision.md
function-catalog.md
scripts/
scan-rounding-calls.R
probe-tie-behavior.R
assets/
report-template.md
evals/
evals.json
fixtures/
myanalysis/
precision-spec.yml
ANSWER-KEY.md
README.md
LICENSE
Each addition has a maintenance purpose:
| Resource | Add it when | Maintenance value |
|---|---|---|
references/ |
Rules or catalogs are substantial, change independently, or apply only to part of the workflow | Keeps each approved rule in one place and lets the entry point load only relevant detail |
scripts/ |
A repeated scan or probe needs deterministic, reviewable behavior | Reduces run-to-run variation and gives technical owners code they can test directly |
assets/ |
Reviewers need a stable deliverable structure | Makes required fields and human gates visible without relying on generated wording |
evals/ |
The skill needs regression evidence across seeded successes, failures, exclusions, and blind spots | Separates test inputs from an independently maintained answer key and makes changes comparable |
README.md and LICENSE |
The package is shared beyond its authors | States intended use, requirements, installation context, limitations, and reuse terms |
This structure is a worked example, not a universal checklist. Empty folders and copied boilerplate create maintenance work without improving the workflow. Conversely, leaving a changing rule, repeated scanner, or benchmark corpus inside one long instruction file makes ownership and testing harder.
7.3.4 Define promotion signals
Decide during Design what evidence would justify moving beyond one file:
- repeated manual probe logic is a signal for a deterministic script;
- long, conditional, or independently governed rules are a signal for focused references;
- inconsistent report sections are a signal for an output template;
- missed candidates, false findings, or unstable runs are signals for fixtures and regression evaluations;
- multiple contributors or consumers are signals for documented ownership, versioning, release notes, and distribution controls; and
- routine operational use is a signal for access, monitoring, incident, rollback, and retirement procedures outside the skill folder.
These are promotion signals, not automatic actions. The accountable owner still decides whether the expected benefit justifies the additional component and its maintenance burden.
7.3.5 Canonical artifact: Task contract
- Objective: Report candidate divergences from BR-001 in one bounded R source set.
- Trigger: An approved on-demand request; automated triggers remain future options.
- Required context: Source files or pasted snippets, target commit when supplied, BR-001 v1.0, the function catalog, required display precision, and the call’s intended report use.
- Minimum checks: Inventory the source, search a catalog of possible rounding and numeric-to-text functions, then probe each provisionally in-scope numeric call with ties of both signs at its required precision and with negative values that should display as zero.
- Agent work: Search beyond the catalog for wrappers and other unexpected paths, provisionally determine whether each candidate affects a displayed numeric value or a report decision, explain any divergence, and draft the report.
- Required output: A Markdown report with scope, coverage, candidates, probe evidence, draft findings, unresolved items, unchecked rules, and human decisions. Catalog and exploratory candidates remain labeled separately.
- Boundaries: Read and test only; no source edits, policy changes, classification approval, or external posting.
- Stop conditions: Missing rule version, unreadable files, unresolved display precision or tie intent, a documented conflicting convention, or failed checks.
- Human decision: The rule owner confirms triage, classifies each finding, approves exceptions, and authorizes any communication.
The contract describes the target behavior. A one-file skill can require the scan and probe, but prose alone cannot make every run deterministic. That limitation is recorded for Test rather than hidden behind extra machinery.
The Design gate is approval of the rule scope, task contract, initial package shape, promotion signals, and stop conditions by the rule owner and an R programmer. Acceptance authorizes a prototype, not a finding or a code change, and initiates Build.
7.4 Build — Prototype the workflow
7.4.1 Level 1: Build the one-file prototype
The book prototype is the one-file skill at rounding/SKILL.md. The complete package contains no supporting files:
rounding/
SKILL.md
The file has five responsibilities:
| Section | Purpose |
|---|---|
| Frontmatter | Names the skill and defines when it applies |
| Rule and scope | Carries BR-001 and identifies relevant report uses |
| Review procedure | Establishes the minimum search, context review, and probe |
| Required report | Keeps coverage, evidence, judgment, and human decisions visible |
| Boundaries | Stops source changes, policy decisions, and external posting |
The description is deliberately narrow. It should match a request to review R rounding under BR-001 without attracting general R review or work on BR-002 and BR-003. The body gives the agent enough direction to perform the review while leaving room to follow wrappers or unexpected code paths.
The skill uses a direct Rscript -e probe rather than a helper script. This keeps the lab self-contained and makes the executed expression visible in the report. The tradeoff is important: prose cannot enforce a complete syntax scan or a machine-readable output schema. Those remain possible future increments if evaluation shows that the one-file skill misses source or produces inconsistent reports.
7.4.2 Level 2: Separate components that need independent maintenance
The RConsortium/pharma-skills/rounding example shows the next build level. Its SKILL.md remains the task router and human-boundary statement, while supporting files take on work that benefits from separate ownership and verification:
- three rule references keep tie method, rounding stage, and display precision distinct;
- a function catalog documents the deterministic scan’s minimum surface and known blind spots;
- one R script scans parse trees and another produces executable witnesses;
- a report template stabilizes the review handoff; and
- a synthetic package, precision specification, evaluation cases, and answer key provide regression material.
This separation helps a team review a rule change without treating it as a scanner change, test a scanner change without rewriting the workflow prose, and revise a report contract without hiding the difference in a prompt. The entry point links to each resource at the step where it is needed, preserving progressive disclosure.
The community package covers BR-001, BR-002, and BR-003, while the book’s one-file prototype intentionally covers only BR-001. Expanding rule scope is a new approved requirement, not a mechanical refactor. A team should first prove the one-rule behavior, then add each new rule with its own owner, reference, test cases, and acceptance criteria.
7.4.3 Level 3: Add the operating system around the package
Production usage needs controls that do not belong entirely inside SKILL.md: an approved version, controlled installation, dependency and permission review, an accepted benchmark in the target environment, named support ownership, monitoring, incident handling, rollback, and retirement. The skill package is the governed unit being released; it is not the complete production system.
7.4.4 Preserve the decisions that made the check useful
The agent may explore beyond the named function catalog, but it must preserve the minimum coverage record. Every candidate remains visible as in scope, excluded, or unresolved. Search matches become draft findings only after source context and an executed probe support the divergence.
The output also keeps three decisions separate:
- BR-001 states the illustrative policy.
- The agent judges whether the evidence indicates a candidate divergence.
- The accountable human decides whether the divergence is a defect, accepted convention, exception, or not applicable.
7.4.5 Canonical artifact: Prototype
For the book lab, the artifact is the versioned SKILL.md itself. For a maintainable package, the artifact is the versioned folder plus a traceable record of its rule, script, template, and evaluation versions. The Build gate requires the rule owner and an R programmer to confirm that:
- BR-001 and its scope match the approved exercise;
- the procedure tests ties of both signs and negative-zero display;
- candidates are not presented as findings without source and probe evidence;
- coverage and unchecked rules are required; and
- the agent stops before decisions reserved for a person.
For the comprehensive level, the gate also requires every supporting resource to have a named purpose, an owner, a link from the entry point or evaluation workflow, and a relevant check. Files that cannot meet those conditions should not be added merely to make the package look mature.
Acceptance makes the skill ready for the small evaluation in Chapter 8 and initiates Test.
7.5 Test — Benchmark and evaluate
The Test stage asks whether the skill improves the review, not whether the folder looks complete. Deterministic scripts can establish what was scanned and executed. They cannot establish that the agent chose the right scope, interpreted the evidence correctly, or produced a useful handoff.
7.5.1 Grow evaluation with the skill
The one-file lab in Chapter 8 uses one synthetic snippet and a small answer key. That is enough to reveal whether the instruction finds the two seeded BR-001 problems, records coverage, and stops at the human decision.
The comprehensive community package demonstrates the next layer in evals/: a synthetic R package, an approved precision specification, evaluation prompts, and an answer key kept out of the run input. The fixture includes direct rounding, early rounding, lost trailing zeros, a compliant helper, allowlisted behavior, an operator without a function name, a dynamic call the scanner cannot see, and out-of-scope negative controls.
Evaluation should grow in layers with the implementation:
| Layer | What it evaluates | Why it matters |
|---|---|---|
| Format validation | Frontmatter, links, parseability, and package layout | Finds broken packaging before behavior is tested |
| Deterministic component checks | Scanner inventories and probe results | Finds script regressions without depending on model judgment |
| Behavioral evaluations | Selection, coverage, verdicts, false findings, blind spots, report completeness, and repeated-run agreement | Tests the assembled skill rather than isolated files |
| Independent human review | Answer-key comparison, clinical relevance, and disposition | Keeps the skill from grading its own interpretation |
| Operational acceptance | Target environment, permissions, latency, cost, observability, rollback, and support response | Tests suitability for the organization’s intended use |
A maintained benchmark should include compliant and noncompliant programs, catalog and uncataloged calls, justified exclusions, missing precision, and a case that the deterministic scanner cannot see. It should record selection accuracy, detected and missed seeded problems, unsupported findings, file and candidate coverage, report completeness, repeated-run agreement, review time, and human triage agreement.
7.5.2 Write the eval before the skill instructions
The layers above describe what to measure. They do not say when to write the eval, and the order matters. Anthropic’s skill-authoring guidance likewise recommends starting with evaluation: run representative tasks to find capability gaps, then build skills incrementally to address them. Write the benchmark rows first, run the base agent without the skill, and record the failing run before writing the skill instructions. The failing run (red) shows which seeded problems the unaided agent misses and which unsupported findings it produces. The skill instructions are then written to close those observed gaps, and the same benchmark is re-run (green) to confirm the improvement.
Concretely, for the next skill built this way:
- Publish the eval first: the benchmark rows and the scoring rule, with no skill instructions present. Keep the answer key outside the evaluated agent’s accessible environment (a sealed location the agent cannot search, not the same repository or network path) and reveal it only after scoring, so neither the red nor the green run can retrieve the expected labels in advance.
- Run the base agent against the eval and record the red result: detected and missed seeded problems, unsupported findings, and coverage.
- Write the skill instructions to address the observed misses, keeping the eval and answer key unchanged.
- Re-run the same eval and record the green result next to the red one, so the improvement is comparable across identical inputs.
Fixing the benchmark before the instructions exist keeps the skill from grading its own interpretation: a passing run means the instructions closed a demonstrated gap rather than restating what the eval already checks. The identical red-green eval measures fit to the development set, so it cannot serve as acceptance evidence on its own: a green result advances only with held-out cases the skill was not tuned against, or a fresh independently authored benchmark run before the Test gate accepts it.
7.5.3 Use pinned source evidence without calling it skill performance
The following evidence remains a useful benchmark seed. All locations refer to metalite.ae v0.1.4 at commit bdb23d472b16bc9dadbc774e64c5ca40321e9c6b. It was manually reverified at that commit; it is not output from either skill.
The review accounted for 27 R files and narrowed the evidence in three steps:
- A search for
round(found one call atR/format_ae_specific.R:330, where a rounded percentage controls row inclusion. - A seven-function catalog scan found 21 candidates: one
round(), eightformatC(), onesprintf(), and 11as.character()calls. - Context review excluded 13 character, structural, or unrelated calls, leaving eight in scope: the row-inclusion decision and seven numeric
formatC()display calls.
The numeric display calls occur at R/fmt.R:33, 75, 79, 99, 100, and 122 and R/format_ae_exp_adj.R:208. The catalog also preserved excluded candidates. For example, R/fmt.R:37 pads text that was formatted at line 33; it does not round numeric input. This correction to issue #249 shows why scanner hits are coverage evidence, not automatic findings.
A one-decimal probe run with R 4.6.1 on arm64 macOS found divergence from the illustrative BR-001 result in 10 of 10 round() values and 6 of 10 formatC() values. The compact zero-decimal checks expose the core behavior:
| Check | Observed | BR-001 |
|---|---|---|
round(2.5, 0) |
2 |
3 |
round(-2.5, 0) |
-2 |
-3 |
formatC(-0.5, digits = 0, format = "f") |
"-0" |
"-1" |
A companion probe, formatC(c(-0.04, 0.04), digits = 0, format = "f"), returned "-0" and "0" under R 4.5.3: a displayed negative zero can also arise when a nonzero value rounds to zero, not only at a tie. BR-001 forbids the display in both cases; only the tie case tests the half-away direction.
Two causes must remain separate. The Base R documentation describes round() as round-to-nearest-even at a true tie. At nonzero decimal precision, binary representation can place the stored value just above or below the intended tie. Source inspection alone cannot predict every direction, so the executed probe is evidence and the explanation is interpretation.
These counts and probes establish an answer key for the pinned source. They do not prove that every wrapper or dynamic path was found, that BR-001 governs the package, or that an agent reproduces the review.
7.5.4 Set acceptance criteria for each level
| Concern | One-file prototype | Maintainable package | Production usage |
|---|---|---|---|
| Selection | Applicable and out-of-scope prompts | Broader selection suite after every description change | Monitored false activation and missed activation in the target setting |
| Coverage | Named snippets or files and unchecked rules | Full fixture inventory, exclusions, wrapper/operator cases, and scanner blind spots | Coverage trend and documented unreviewed surfaces |
| Accuracy | Seeded findings and no false finding in the small answer key | Rule-level PASS, FAIL, and NOT ASSESSABLE cases across the maintained corpus | Accepted target-environment thresholds and independent review |
| Evidence | Visible direct probe | Versioned scripts plus before-and-after witnesses | Reproducible run records with environment and package versions |
| Boundaries | No source edits, policy decisions, or external posting | Permission and issue-mode tests | Enforced access, authorization, incident, and rollback controls |
| Reproducibility | Repeat the same prompt and compare | Repeated runs across supported models or harnesses | Release-specific monitoring for drift, latency, cost, and failures |
7.5.5 Canonical artifact: Benchmark report
No formal benchmark report exists for the book’s one-file prototype. Chapter 8 provides a guided observation, while the community package provides a broader corpus. The presence of fixtures and an answer key is not itself an accepted result for a particular model, host, or organization.
A benchmark report must identify the skill version, target, environment, rules, and precision; compare every run with an independently maintained answer key; separate deterministic output from agent judgment; report misses, false findings, coverage, stopped runs, reproducibility, time, and limitations; and record the reviewers’ disposition.
The Test gate remains open for the book prototype. At either package level, the rule owner and an independent reviewer must accept the executed benchmark report. A production candidate also needs results in its target model, harness, R environment, and permission configuration. Only that accepted evidence would initiate Deploy.
7.6 Deploy — Operationalize the workflow
7.6.1 Deploy the version that was tested
The first deployment should remain modest: an on-demand, repository-scoped advisory review. The skill may read the named source, run small isolated probes, and write a draft report. It receives no permission to modify source or communicate externally.
The lab references rounding/SKILL.md directly. This keeps the source artifact easy to find but does not make it an auto-discovered repository skill. For a later Codex installation, copy the rounding directory under .agents/skills/ in the working directory or repository where it should be available. The installed path becomes .agents/skills/rounding/SKILL.md.
A maintained package adds distribution choices, but the rule is the same: install an immutable version that matches the accepted benchmark, record where it is enabled, and preserve a way to disable or roll it back. Do not test one commit and deploy a moving main branch. Review executable scripts and their permissions before installation, and keep external writes disabled unless a separate task contract and approval explicitly allow them.
| Deployment scope | Minimum release evidence |
|---|---|
| Guided one-file trial | Reviewed SKILL.md, small answer key, named reviewer, and bounded read-only use |
| Shared maintainable package | Immutable package version, accepted regression report, dependency and permission review, documented installation, owner, and rollback |
| Routine production use | Target-environment acceptance, controlled change and distribution, monitoring thresholds, support and incident process, periodic review, and retirement criteria |
The RConsortium/pharma-skills/rounding directory is a community distribution example. Its README.md, LICENSE, versioned skill metadata, and repository lifecycle make shared maintenance more visible than the book lab. They do not by themselves authorize installation or establish fitness for an organization’s production use.
Issue #249 was posted manually from an exploratory response outside this workflow. It is not a Deploy artifact or evidence that the skill was released or that BR-001 governs metalite.ae.
7.6.2 Canonical artifact: Release record
- Book status: Not released; design-pattern evaluation only.
- Book candidate: The one-file
roundingskill. - Comprehensive case study: Community version 0.9 at the pinned commit; inspect its current evidence and lifecycle before treating a later version as equivalent.
- Release model: Manual and advisory; a reviewer requests one run.
- Scope: One bounded R source set and BR-001 version 1.0.
- Permissions: Read source, run isolated local probes, and draft Markdown; no source or network writes.
- Accountable human: The named rule owner approves release, classifications, exceptions, and any external communication.
- Reversal: Remove or disable the skill. A normal advisory run changes no source or external system.
Scheduling, automated change detection, and issue publishing are not part of the book release. If later evidence justifies them, they require their own requirements, risks, tests, permissions, and approval gates. A capability present in a comprehensive package does not become authorized merely because the package was installed.
Acceptance of the release record after the benchmark would initiate Maintain. Neither event has occurred in this book.
7.7 Maintain — Monitor and improve
7.7.1 Make improvement a governed change
An operating check could suggest catalog additions, clearer instructions, or rule clarifications. Those are proposals. Before a change becomes active, the rule owner approves policy changes, an R programmer reviews technical changes, and an independent reviewer reruns the affected test examples. The agent does not edit the rule against which it is evaluated.
The comprehensive structure makes impact analysis possible when ownership is explicit:
| Changed concern | Primary artifact | Required regression scope |
|---|---|---|
| Business rule or approved helper | Matching references/br-00x-*.md and rule version |
Rule-specific positive, negative, boundary, and unresolved cases |
| Candidate discovery | Scanner and function catalog | File inventory, wrapper/operator cases, exclusions, and known blind spots |
| Executed evidence | Probe script | Supported R environments and every required precision |
| Review handoff | Report template | Required fields, evidence separation, and human decision gate |
| Activation boundary or procedure | SKILL.md |
Applicable and out-of-scope requests plus end-to-end fixture runs |
| Distribution or dependency | Package metadata and release record | Installation, permissions, provenance, rollback, and compatibility |
SKILL.md remains the map across these components. A change is incomplete when the entry point routes to an obsolete reference, a script changes without its answer key, or a new rule appears without a responsible owner. This is the practical maintenance benefit of modularity; folder depth by itself provides none.
The first list of improvements for the one-file version follows directly from known limitations:
- Add a rule-owner-approved allowlist for compliant wrappers.
- Record target commit, versions, files scanned, exclusions, probe output, triage reasons, and classifications in a run log.
- Report files or expressions the check could not evaluate.
- Keep this BR-001 skill bounded. Use the implemented three-rule
RConsortium/pharma-skills/roundingpackage as a design input, but approve BR-002 and BR-003 separately before expanding the book skill’s active scope.
7.7.2 Canonical artifact: Monitoring report
There is no operating history because the skill has not been released. The monitoring report is therefore specified, not populated. A future report must show applicable requests that did and did not select the skill, unrelated requests that selected it incorrectly, runs by trigger, coverage, confirmed findings, false positives, missed defects, stopped runs, triage effort, and unresolved precision. It also records how long an approved BR-001 change takes to reach the skill and its tests, whether the active skill still names the current rule version, incidents, and approved changes.
The rule owner reviews that report and chooses among continued use, correction, promotion to a more maintainable package, retirement, or a new planning cycle. An accepted policy change, incident, measured coverage gap, new operating scope, or dependency change is the event that starts the next cycle.
7.8 Summary
Chapter 6 showed that a written rule and one open prompt can produce useful candidates, examples, explanations, and a proposed solution. It did not treat that response as an answer key. The separate source review in this chapter identifies one round() call used in a report-filtering decision and seven in-scope numeric formatC() calls. Its one-decimal probe diverges from BR-001 in 10 of 10 round() results and 6 of 10 formatC() results in the recorded environment. These are differences from the illustrative rule, not a maintainer-approved defect classification. Tie mode and binary representation remain separate causes.
The distributed book skill packages the trigger, BR-001, procedure, report, and boundaries in one SKILL.md. It may be proportionate for a small, low-risk review, but one file does not guarantee reliable behavior. A future benchmark must evaluate selection, findings, false findings, coverage, output completeness, and repeated-run agreement.
Chapter 8 keeps the first evaluation equally small: one synthetic source, one invocation, one answer key, and one human review. An accepted benchmark report, release, and operating evidence remain future work.
The RConsortium/pharma-skills/rounding package demonstrates the maintainable next level: SKILL.md remains the entry point while rule references, deterministic R scripts, a report asset, and evaluation fixtures can change under focused review. It covers BR-001, BR-002, and BR-003 and grew from the benchmark discussion in RConsortium/pharma-skills issue #208.
The strategic path is therefore one bounded contract -> evidence-backed modules -> controlled production operation. Promotion occurs only when the current level exposes a measured gap and the next component has a purpose, owner, test, and release gate. That is how a skill becomes maintainable and ready to be evaluated for production use without confusing package complexity with proven fitness.