Repository navigation
feat(extract,detect): Apps Script, CSS/SCSS design tokens and an explicit secrets-filter override (#2498, #2698, #3473) - #4263
Conversation
Apps Script sources are plain JavaScript saved as `.gs` — what `clasp` pulls down and what every Apps Script project in a repo looks like. `.gs` was absent from `CODE_EXTENSIONS`, the extractor `_DISPATCH`, and the JS language-family maps, so an entire Apps Script project was classified as non-code and contributed nothing: a build over a repo whose logic lives in `.gs` returned only its markdown and `appsscript.json` manifests. On a 29-file Apps Script project this was 83 nodes / 84 edges before and 285 / 835 after, with every god node before the change coming from a doc. `.gs` is not exclusive to Apps Script — GLSL geometry shaders and Gosu use it too — so routing goes through `_is_apps_script`, which withholds the extractor when a GLSL or Gosu marker is present rather than force-parsing the file into garbage. The sniff runs the opposite way from the `.m` Objective-C/MATLAB split: `.m` has a genuine rival and is guilty until proven innocent, while `.gs` in a repo is nearly always Apps Script and its own markers are unusable as a positive test (a pure-logic `.gs` helper calls no `SpreadsheetApp` or `DriveApp` service at all). `_JS_TS_CALL_SUFFIXES` deliberately does not gain `.gs`: that gate drops cross-file calls lacking import evidence, which is right for ES modules and wrong for Apps Script, where every file shares one global scope and imports do not exist.
…s/.css/.scss/.less and dotfile configs (Graphify-Labs#2498, Graphify-Labs#2698, Graphify-Labs#3473) An explicit wildcard-free '!path' entry in .graphifyignore now overrides only the Stage-3 secrets-keyword heuristic, never the .env/key-file or credential-store stages; a rescued file is kept and the skip report names the escape hatch (Graphify-Labs#2498). Well-known extensionless config/build files (.htaccess, Dockerfile, Makefile, ...) are classified as documents instead of dropped; .less routes to the CSS extractor (Graphify-Labs#2698). Combines and supersedes Graphify-Labs#2747/Graphify-Labs#3505 (.gs Apps Script), Graphify-Labs#3480/Graphify-Labs#3067/Graphify-Labs#2240 (CSS/SCSS design tokens) and Graphify-Labs#1921 (allowlist) against current v8.
…t ignores; stop false CSS token edges (Graphify-Labs#2498, Graphify-Labs#3473) Verification follow-up. The rescue now consults only the explicit .graphifyignore/--exclude pattern set (never .gitignore), is anchored to the entry's own directory so a nested ignore cannot rescue a same-named file elsewhere, and matches case-sensitively. A custom property declared outside a root/theme context now marks its stylesheet's var() uses as locally shadowed, so cross-file resolution no longer invents a false EXTRACTED uses_token edge.
… override entries (Graphify-Labs#2498, Graphify-Labs#3473) A stylesheet that defines a token in :root and overrides it locally still gets its same-file uses_token edge; the local shadow now only suppresses cross-file binding. A leading-slash .graphifyignore entry (!/name) is anchored to its own directory again, matching gitignore semantics.
|
Thanks for the pull request, @nothariharan. A maintainer will review it soon. Want to talk it through while it is in review? Come join us on our Discord server. For longer-form discussion there is also GitHub Discussions. A couple of things that speed up review: make sure the test suite passes on Python 3.10 and 3.13, and that the change keeps extraction deterministic. |
There was a problem hiding this comment.
Graphify reviewed this change.
Code graph out of date: Health delta baseline was built by an older engine whose node ids do not match this review's — delta is approximate. Impact and hotspot numbers below may be wrong.
Formal verification could not match the changed code to the code graph, so it may not have checked the code this PR changed.
No verdict: the code graph is out of date. Static checks found no blocking issues, but they compared against an outdated graph. The next push after the repository re-indexes gets a full review.
Not checked: tests were not run; formal verification proved 0 of 9 changed function(s) (3 sampled, not proven, 6 not verified).
Formal verification. PR-changed functions: 3/9 verified (0 proven, 3 may-equivalent, 0 distinguished) · 6 not verified (3 vacuous, 3 unsupported).
Not verified on this run: \_absolutize\_source\_files\_in (unsupported), \_relativize\_source\_files\_in (unsupported), dispatch\_command (vacuous: never exercised), detect (unsupported), extract (vacuous: never exercised), generate (vacuous: never exercised).
Graphify review — findings
Adds Google Apps Script .gs files as JS-family code, and also classifies .css/.scss/.less as code. Lets a wildcard-free !<path> entry in .graphifyignore or --exclude rescue a file dropped by the name-keyword sensitive heuristic; credential-store dirs and specific secret patterns like .env or id_rsa are never overridable, and the skip warning now explains the escape hatch. Cache entries now relativize and re-absolutize raw_token_uses paths along with nodes, edges, and calls, so warm hits no longer replay the build host's absolute layout.
Review partial — this diff was larger than one review pass covers, so later files were not reviewed; some findings may be missing.
Analysis details — impact, health, verification
Impact & health
Graphify review
Impact — 4926 functions depend on the 1325 functions this change touches.
Health — this change adds coupling hotspots:
- new:
extract()— 829 callers, 52 callees - new:
_rebuild_code()— 162 callers, 56 callees - new:
build_from_json()— 235 callers, 21 callees - new:
detect()— 129 callers, 17 callees - new:
build_merge()— 85 callers, 14 callees - new:
save_manifest()— 48 callers, 15 callees - new:
_extract_generic()— 19 callers, 34 callees - new:
save_semantic_cache()— 65 callers, 9 callees - …and 134 more — each is listed as a finding
Verification — 4926 functions in the blast radius were not formally verified this run (proofs are advisory here).
Health delta baseline was built by an older engine whose node ids do not match this review's — delta is approximate.
Gate & verification
graphify gate
PASS — objectively clean (no health regressions, tests not run — proofs not run this pass (advisory)). Grounded, not self-assessed.
Advisory (not blocking):
- verification_scope: 4669 function(s) in the blast radius were not formally verified this run
Test selection
Test selection
357 of 357 test file(s) selected (100%) via static blast radius.
Escalated to a full run for safety — the selection is not trustworthy on its own (see below). CI should run the whole suite.
tests/test_affected_cli.py— impact, full-run-safetytests/test_affected_member_seed.py— full-run-safetytests/test_agents_platform.py— impact, full-run-safetytests/test_analyze.py— impact, full-run-safetytests/test_anthropic_custom_endpoint.py— full-run-safetytests/test_antigravity_install.py— full-run-safetytests/test_apm_fallback_version.py— full-run-safetytests/test_architecture_doc.py— full-run-safetytests/test_astro_extraction.py— impact, full-run-safetytests/test_astro_import_ids.py— impact, full-run-safetytests/test_atomic_canvas_export.py— impact, full-run-safetytests/test_atomic_version_stamp.py— full-run-safetytests/test_atomic_writes.py— impact, full-run-safetytests/test_backend_env_isolation.py— full-run-safetytests/test_backend_extras.py— full-run-safetytests/test_benchmark.py— impact, full-run-safetytests/test_benchmark_raw_graph.py— impact, full-run-safetytests/test_blade_extractor.py— impact, full-run-safetytests/test_build.py— impact, full-run-safetytests/test_build_located_semantic_identity.py— impact, full-run-safetytests/test_build_merge_dedup_scope.py— impact, full-run-safetytests/test_build_merge_hyperedges_and_prune.py— impact, full-run-safetytests/test_build_merge_shrink_guard.py— impact, full-run-safetytests/test_builtin_global_type_refs.py— impact, full-run-safetytests/test_cache.py— impact, full-run-safetytests/test_cache_stale_import_target.py— impact, full-run-safetytests/test_callflow_html.py— full-run-safetytests/test_cargo_introspect.py— impact, full-run-safetytests/test_cargo_missing_manifest.py— full-run-safetytests/test_carried_hyperedge_remap.py— impact, full-run-safetytests/test_case_sensitive_resolution.py— impact, full-run-safetytests/test_charmap_encoding.py— impact, full-run-safetytests/test_chunking.py— impact, full-run-safetytests/test_cjs_module_extension.py— impact, full-run-safetytests/test_claude_cli_backend.py— impact, full-run-safetytests/test_claude_md.py— full-run-safetytests/test_cli_broken_pipe.py— full-run-safetytests/test_cli_export.py— impact, full-run-safetytests/test_cli_help.py— full-run-safetytests/test_cloud_cta.py— impact, full-run-safetytests/test_cluster.py— impact, full-run-safetytests/test_cluster_ambiguous_scale.py— full-run-safetytests/test_cluster_exclude_hubs.py— full-run-safetytests/test_cobol_extractor.py— impact, full-run-safetytests/test_codebuddy.py— impact, full-run-safetytests/test_community_hub_labels.py— full-run-safetytests/test_community_labels_skill.py— impact, full-run-safetytests/test_confidence.py— impact, full-run-safetytests/test_corrupt_graph_json.py— impact, full-run-safetytests/test_cpp_method_declarations.py— impact, full-run-safety- … and 307 more
non-code file(s) changed (
README.md) → running the full suite for safety (a code graph can't see config/fixture/data deps)
changed code file(s) with no mapped test (
README.md,graphify/extractors/__init__.py) — a coverage gap or a missing link — running the full suite rather than only the selected tests
Selection is safe under the controlled-regression assumption; always-run tests + a periodic full run are the backstops. Advisory — it never changes the check verdict.
Docs that may be stale (advisory)
CHANGELOG.md§ 0.9.18 (2026-07-17) (lines 805-818): references changed symbols_is_sensitiveCHANGELOG.md§ 0.9.8 (2026-07-06) (lines 984-994): references changed symbols_is_sensitiveCHANGELOG.md§ 0.8.34 (2026-06-07) (lines 1317-1334): references changed symbols_is_sensitiveCHANGELOG.md§ 0.8.12 (2026-05-18) (lines 1511-1521): references changed symbols_is_sensitiveCHANGELOG.md§ 0.7.6 (2026-05-05) (lines 1729-1741): references changed symbols_is_sensitive
Formal verification
Could not verify: Could not verify \_absolutize\_source\_files\_in.
The verifier did not have enough to check \_absolutize\_source\_files\_in, so it is saying so rather than guessing. No false assurance is the whole point.
Guarantee: No guarantee either way, this is an honest abstention, not a pass.
Note: Reason: no capturable inputs from the test suite; property tier: not verifiable: the input domain has 59 values but only 24 distinct were tested — a small finite domain must be EXHAUSTED, not sampled (an untested input could invert the result)
Could not verify: Could not verify \_relativize\_source\_files\_in.
The verifier did not have enough to check \_relativize\_source\_files\_in, so it is saying so rather than guessing. No false assurance is the whole point.
Guarantee: No guarantee either way, this is an honest abstention, not a pass.
Note: Reason: no capturable inputs from the test suite; property tier: not verifiable: the input domain has 59 values but only 24 distinct were tested — a small finite domain must be EXHAUSTED, not sampled (an untested input could invert the result)
Could not verify: Could not verify dispatch\_command.
The verifier did not have enough to check dispatch\_command, so it is saying so rather than guessing. No false assurance is the whole point.
Guarantee: No guarantee either way, this is an honest abstention, not a pass.
Note: Reason: no capturable inputs from the test suite; property tier: not verifiable: all 40 sampled inputs raised on both versions — the function never executed, so 'no divergence' would be vacuous (mostly SystemExit — names the real obstacle, not a sampling gap)
No difference found (not proven): No behavior difference found in \_is\_sensitive (not a proof).
The verifier ran both versions of \_is\_sensitive on many inputs and saw identical behavior every time. Strong evidence the change is safe, but evidence, not a proof.
Guarantee: Empirical: differential testing (both versions run on many generated inputs). A divergence on an untested input remains possible, so this is 'no counterexample found', not 'proven equivalent'.
Note: An input the sampler did not try could still differ.
No difference found (not proven): No behavior difference found in classify\_file (not a proof).
The verifier ran both versions of classify\_file on many inputs and saw identical behavior every time. Strong evidence the change is safe, but evidence, not a proof.
Guarantee: Empirical: differential testing (both versions run on many generated inputs). A divergence on an untested input remains possible, so this is 'no counterexample found', not 'proven equivalent'.
Note: An input the sampler did not try could still differ.
Could not verify: Could not verify detect.
The verifier did not have enough to check detect, so it is saying so rather than guessing. No false assurance is the whole point.
Guarantee: No guarantee either way, this is an honest abstention, not a pass.
Note: Reason: no capturable inputs from the test suite; property tier: nondeterministic output (both versions disagreed with themselves on the witness) — not a behaviour change; a function whose output is not a function of its inputs (random / time / uuid / os.urandom, or state carried across calls) cannot be soundly distinguished by execution, so this tier abstains for the whole function
No difference found (not proven): No behavior difference found in \_get\_extractor (not a proof).
The verifier ran both versions of \_get\_extractor on many inputs and saw identical behavior every time. Strong evidence the change is safe, but evidence, not a proof.
Guarantee: Empirical: differential testing (both versions run on many generated inputs). A divergence on an untested input remains possible, so this is 'no counterexample found', not 'proven equivalent'.
Note: An input the sampler did not try could still differ.
Could not verify: Could not verify extract.
The verifier did not have enough to check extract, so it is saying so rather than guessing. No false assurance is the whole point.
Guarantee: No guarantee either way, this is an honest abstention, not a pass.
Note: Reason: no capturable inputs from the test suite; property tier: not verifiable: all 90 sampled inputs raised on both versions — the function never executed, so 'no divergence' would be vacuous (mostly TypeError — names the real obstacle, not a sampling gap)
Could not verify: Could not verify generate.
The verifier did not have enough to check generate, so it is saying so rather than guessing. No false assurance is the whole point.
Guarantee: No guarantee either way, this is an honest abstention, not a pass.
Note: Reason: no capturable inputs from the test suite; property tier: not verifiable: all 264 sampled inputs raised on both versions — the function never executed, so 'no divergence' would be vacuous (mostly ValueError — names the real obstacle, not a sampling gap)
· 3 grounded finding(s) anchored inline below; 139 more finding(s) on lines outside this diff (see the check run).
| }) | ||
|
|
||
|
|
||
| def classify_file(path: Path) -> FileType | None: |
There was a problem hiding this comment.
classify_file()
53 callers depend on it (afferent coupling).
Grounded coupling-delta finding (deterministic), not an LLM guess.
| return not any(marker in head for marker in _NON_APPS_SCRIPT_GS_MARKERS) | ||
|
|
||
|
|
||
| def _get_extractor(path: Path) -> Any | None: |
There was a problem hiding this comment.
_get_extractor()
fans out to 7 callees (efferent coupling); 32 callers depend on it (afferent coupling).
Grounded coupling-delta finding (deterministic), not an LLM guess.
| return True | ||
|
|
||
|
|
||
| def extract_css(path: Path) -> dict: |
There was a problem hiding this comment.
extract_css()
13 callers depend on it (afferent coupling).
Grounded coupling-delta finding (deterministic), not an LLM guess.
Consolidates the outstanding fixes for three issues into one branch, rebased onto current
v8. Every referenced PR was raised against an olderv8revision and no longer applies; this is the single, current version. One previous version of this PR was self-reviewed and two real issues were found and fixed before opening (see Review notes below).What this fixes
#2698 — classify
.gs,.cssand extensionless config files.gs(Google Apps Script) is now code: routed to the JavaScript extractor, the JS language/resolution family, the AST-cache bypass set and the rebuild hook. A GLSL/Gosu marker guard (#version,gl_Position,EmitVertex,uses java.…) leaves a real geometry shader / Gosu file without an extractor rather than force-parsing it as JS..css/.scss/.lessare code..htaccess,Dockerfile,Makefile,Procfile,Vagrantfile, …) are classified by filename as documents instead of being dropped as "not classified".#3473 — design-token systems were invisible to the graph
graphify/extractors/css.py: a brace-aware, comment/string-aware, dependency-free scan. It emits a node per custom-property definition in a root/theme context (:root,:host,html,body,[data-theme…],.dark,.theme-*,@theme), adefines_tokenedge per definition, and collectsvar(--x)uses.:rootkeeps its own same-file edge.raw_token_uses) and both id-remap passes; warm and cold extraction agree.#2498 — the filename-only secrets filter silently drops design-token files
The report has two distinct halves:
.csscase (reporting defect):tokens.cssis now genuinely graphable, so the misleading "skipped as potentially sensitive" message for an unsupported extension is gone.tokens.json/tokens.contract.json(real data loss):#3527deliberately keeps bare pluraltokensin data files excluded, so rather than weaken the default the fix is the explicit override the issue asked for (option 3): a wildcard-free!pathentry in.graphifyignore(or--exclude) now overrides only the Stage-3 name heuristic..env,id_rsa,server.pem) or a credential-store directory is never rescued.!*.json) are ignored, so a broad opt-in cannot silently start ingesting credential stores./is anchor-exact, per gitignore semantics..graphifyignore/--excludeset is consulted — never.gitignore.tokens.json,oauth_token.json,api_token.txt,credentials.json,.token,secrets.jsonall stay skipped unless explicitly overridden.GRAPH_REPORT.mdgains a## Skipped Filessection listing every excluded file by name, and the CLI skip notice names the escape hatch.PRs this supersedes
.gsonly, minimal.gssupport here.gs(comprehensive).graphifyincludeoverride.graphifyignorenegation (.graphifyincludewas removed in #2112)All of the above no longer merge onto
v8; this branch applies cleanly onv8@5b74d7d.Review notes (self-review before opening)
Two independent verification passes were run. They caught and this branch fixes:
.gitignoreand nested anchors) and matching case-insensitively — i.e. broader than the "exact.graphifyignorepath" it promised. It now uses the explicit pattern set only, is anchor-scoped, and case-sensitive.var()resolution initially ignored local custom-property shadowing, inventing falseEXTRACTEDuses_tokenedges between a component override and a theme token. Fixed, then narrowed so a file that also defines the token in:rootkeeps its same-file edge.Tests
New / extended:
tests/test_gs_apps_script.py—.gsrouting, GLSL/Gosu guard, node parity with.js.tests/test_css.py— CSS/SCSS/Less classification, token extraction, comments/strings/urls/nesting/themes, cross-file resolution, same-file precedence, local shadowing, warm cache.tests/test_detect.py— explicit-negation rescue,.gitignore-only negation does not rescue, anchor scoping, leading-slash anchoring, wildcard entries do not rescue, Stage-1/2 never rescued, dotfile classification.tests/test_report.py—## Skipped Filespresent/absent.Local results (
v8@5b74d7d, Python 3.12, Windows):test_detect.py304 passed / 1 skipped,test_css.py25 passed,test_report.py21 passed,test_gs_apps_script.py10 passed,test_extract.py289 passed / 4 skipped.ruff checkon all changed files: clean.v8checkout — they are pre-existing Windows/environment failures (noos.mkfifo, read-only-file semantics, missing home dir under a stripped env, cp1252 unicode, and gemini/hermes install + watch/hook timing). None is in a file this branch touches.Known limitations
.gssniff scans the first 256 KiB for GLSL/Gosu markers; a genuine Apps Script file that literally contains one of those substrings in a string or comment is left without an extractor (same outcome as MATLAB.m). Kept from feat(gs): extract Google Apps Script as JavaScript #3505's conservative design.var()references are resolved..css/.scss/.less, not Sass indented syntax;@import url(...)prose is not resolved.