What validation actually runs¶
The default development policy favors quick feedback. A routine MR is a bounded check, not a claim that every feature works. Run a small reproduction for the code you changed. Broad numerical validation uses a qualified development build and appropriate compute resources; it need not block every implementation merge.
Inventory and counting¶
Snapshot at source 0819b1771c84 (2026-09-26): 933 test files and 15,028
test_* function/method definitions under tests/. Definitions are not collected
pytest cases: parametrization expands them, and dependencies can skip modules.
A full collected count requires a matching native runtime. No native build was
started just to count tests. vibe-basis/tests/, hook tests and website tests are
separate from these totals.
The exhaustive scripts/test_gate/suite_manifest.json classifies those files:
Tier |
Files |
Purpose |
Default CI execution |
|---|---|---|---|
T1 |
12 |
Blocking sentinels and policy/contract checks |
Candidate or explicit full validation |
T2 |
847 |
Broad regression, numerical and reference inventory |
Selected manually; not the default MR gate |
T3 |
75 |
Experimental/research inventory |
Explicit research runs |
T0 |
0 pytest files |
Compile/import and runtime provenance |
Opt-in native MRs and full validation |
In the measured release baseline, T1 produced 294 passed and two skipped: one skip was a whole module, not an executed passing case. The sum of its file execution times was 128.3 seconds; files run concurrently, so this is not wall time. The combined native build/test job took 56.7 minutes. Compilation, setup and producer/consumer work are additional to T1.
Types of validation¶
Type |
Concrete examples |
Does it run chemistry? |
|---|---|---|
Static and repository contracts |
Privacy, syntax compilation, source inventory, changelog pins, basis coverage |
No SCF calculation |
Unit/algebra tests |
Tensor operations, solver residuals, conventions, option validation; mocks around native calls |
Usually no complete calculation |
Small numerical regression |
H2/H/STO-3G RHF/UHF/RKS/UKS and MP2 fixed points in |
Yes, real small SCF/correlation/gradient calculations |
Broad numerical/reference tests |
Molecular correlation, periodic GDF/GAPW/BIPOLE, basis/ECP, cross-code parity |
Often multiple real calculations; may need large memory or an external reference |
Producer/consumer integration |
QC-generated output consumed by the pinned viewer |
Some tests require the native QC runtime |
Native build |
C++/pybind11 build plus dependency setup and import provenance |
Builds the implementation; compilation alone proves no scientific result |
Other build checks |
Sphinx docs; Astro website tests/build/output verification; vibe-basis tests and install lifecycle |
No production calculation campaign |
T1 itself is mixed. Basis/source/changelog files are contracts; derivative guard files contain both mocked/refusal tests and numerical checks. A single function can run several methods, so test count is not calculation count. Large periodic, CC, scaling and reference acceptance remains a separate measured activity with inputs, runtime identity, tolerances and an independent verdict. Green CI alone is not that verdict. Experimental T3 work does not become a routine release gate.
Execution policy¶
Trigger |
What runs |
|---|---|
Python-only, test-only or CI-YAML-only MR |
Privacy, policy tests, Python syntax and T1 selection checks; relevant companion/tooling jobs; no native rebuild |
C++/native build input MR |
Cheap gates by default; set |
Main push |
Privacy, and website verification when applicable |
Release candidate |
Clean native build, T1, pinned viewer contracts, vibe-basis and schema checks, docs |
Manual web pipeline with |
Explicit clean native/T1/viewer validation plus companion gates |
This intentionally defers native compiler errors, Python runtime errors and
Python/native mismatches to focused local work, a development build or release
validation. There is no default C++ compilation gate on an MR. Opt in with
VIBEQC_BUILD_NATIVE=1 for an MR touching native inputs, or use full validation.
Routine docs edits no longer build the entire Sphinx site. MR docs builds are limited to docs configuration/theme/build scripts and CI configuration; candidates, release and web pipelines still build the site. A prose cross-reference can therefore fail later. QVF schema conformance is path-filtered on MRs, and CI-YAML edits no longer exercise the unrelated vibe-basis installation lifecycle. CI edits no longer rebuild the native core merely because the YAML changed. Superseded validation jobs are interruptible; candidate auto-cancellation remains disabled. Builds that are still needed remain serialized by the native resource group.
Commands¶
# Cheap policy checks, no native extension required.
python3 -m unittest discover -s scripts/ci/tests -p 'test_*.py'
python3 scripts/test_gate/run_full_suite.py --wt . --tier T1 --dry-run
# Affected small tests on an already qualified runtime.
.venv/bin/python -m pytest tests/test_<area>.py -q
# Full blocking release test tier, not the entire research inventory.
.venv/bin/python scripts/test_gate/run_full_suite.py --wt . \
--py .venv/bin/python --tier T1 --out <external-evidence>/blocking.jsonl
# Inventory expansion, only on a runtime matching this source tree.
.venv/bin/python -m pytest tests --collect-only -q
See developer test lanes for targeted lane commands. Avoid launching all T2/T3 calculations or a new native build just for routine feedback or counting. The oldest performance entries in the manifest are historical measurements, not current runtime promises.