Repository navigation
Conversation
…y-Labs#4266) Co-Authored-By: Grok Bot <noreply@x.ai>
|
Thanks for the pull request, @Mpasha17. A maintainer will review it soon. Want to talk it through while it is in review? Come join us on our Discord server. For longer-form discussion there is also GitHub Discussions. A couple of things that speed up review: make sure the test suite passes on Python 3.10 and 3.13, and that the change keeps extraction deterministic. |
There was a problem hiding this comment.
Graphify reviewed this change.
Formal verification could not match the changed code to the code graph, so it may not have checked the code this PR changed.
No issues found by static checks — no coupling regressions or blocking issues in the code graph. That is not the same as safe to merge: see what was not checked.
Not checked: tests were not run; formal verification proved 0 of 4 changed function(s) (1 sampled, not proven, 3 not verified).
Formal verification. PR-changed functions: 1/4 verified (0 proven, 1 may-equivalent, 0 distinguished) · 3 not verified (2 vacuous, 1 unsupported).
Not verified on this run: \_resolve\_csharp\_member\_calls (vacuous: never exercised), \_csharp\_method\_receiver\_types (vacuous: never exercised), \_extract\_generic (unsupported).
Graphify review — findings
Resolves C# member calls on var v = R.M<A>() locals by recording each method's declared return type and type-parameter count during extraction. The cross-file pass in _resolve_csharp_member_calls then types v from M's return, substituting a call-site type argument when M returns its own type parameter. When M resolves to zero or several overloads, its return type can't be named, or R is an untyped lowercase local, the call is skipped rather than guessed.
No blocking issues surfaced. 9 lower-confidence candidates did not survive cross-model review.
Analysis details — impact, health, verification
Impact & health
Graphify review
Impact — 2961 functions depend on the 586 functions this change touches.
Health — this change adds coupling hotspots:
- worse:
extract()— 835 callers, 51 callees - worse:
walk()— 1 callers, 71 callees - worse:
_csharp_method_receiver_types()— 1 callers, 6 callees
Verification — 2961 functions in the blast radius were not formally verified this run (proofs are advisory here).
Gate & verification
graphify gate
PASS — objectively clean (no health regressions, tests not run — proofs not run this pass (advisory)). Grounded, not self-assessed.
Advisory (not blocking):
- verification_scope: 2766 function(s) in the blast radius were not formally verified this run
Test selection
Test selection
161 of 359 test file(s) selected (45%) via static blast radius.
tests/test_astro_extraction.py— impacttests/test_astro_import_ids.py— impacttests/test_blade_extractor.py— impacttests/test_build.py— impacttests/test_builtin_global_type_refs.py— impacttests/test_cache.py— impacttests/test_case_sensitive_resolution.py— impacttests/test_cjs_module_extension.py— impacttests/test_cobol_extractor.py— impacttests/test_cpp_method_declarations.py— impacttests/test_cpp_nested_and_cli.py— impacttests/test_cpp_objc_cross_file_calls.py— impacttests/test_cross_extension_reexport_self_cycle.py— impacttests/test_cross_language_call_resolution.py— impacttests/test_cross_repo_external_call_guards.py— impacttests/test_cross_repo_member_calls.py— impacttests/test_csharp_call_site_generic_args.py— impacttests/test_csharp_enum_members.py— impacttests/test_csharp_field_generic_args.py— impacttests/test_csharp_generic_callsites.py— impacttests/test_csharp_interface_dispatch.py— impacttests/test_csharp_member_calls.py— impacttests/test_csharp_member_chain_calls.py— impacttests/test_csharp_member_nodes.py— impacttests/test_csharp_object_creation.py— impacttests/test_csharp_partial_classes.py— impacttests/test_csharp_tuple_type_refs.py— impacttests/test_csharp_type_resolution.py— impacttests/test_csharp_var_call_return_type.py— impact, changed-testtests/test_definition_file_portability.py— impacttests/test_detect.py— impacttests/test_dotnet.py— impacttests/test_duplicate_annotation_edges.py— impacttests/test_elixir_import_resolution.py— impacttests/test_elixir_keyword_def_calls.py— impacttests/test_elixir_qualified_calls.py— impacttests/test_elixir_unqualified_call_scope.py— impacttests/test_erlang_extractor.py— impacttests/test_extract.py— impacttests/test_extract_cache_location.py— impacttests/test_extract_path_memo.py— impacttests/test_extract_php_closures.py— impacttests/test_file_label_disambiguation.py— impacttests/test_file_node_id_spec.py— impacttests/test_forwarding_review_findings.py— impacttests/test_go_builtin_call_targets.py— impacttests/test_go_import_repoint.py— impacttests/test_go_interface_methods.py— impacttests/test_go_qualified_resolution.py— impacttests/test_import_extension_resolution.py— impact- … and 111 more
Selection is safe under the controlled-regression assumption; always-run tests + a periodic full run are the backstops. Advisory — it never changes the check verdict.
Docs that may be stale (advisory)
docs/superpowers/plans/2026-05-04-incremental-updates-dedup.md§ Task 4: Incremental updates — semantic cache + manifest in __main__.py (lines 603-911): references changed symbols_run
Formal verification
Could not verify: Could not verify \_resolve\_csharp\_member\_calls.
The verifier did not have enough to check \_resolve\_csharp\_member\_calls, so it is saying so rather than guessing. No false assurance is the whole point.
Guarantee: No guarantee either way, this is an honest abstention, not a pass.
Note: Reason: no capturable inputs from the test suite; property tier: non-vacuity: domain too small (only 1 distinct inputs exercised, need 3) — 'no divergence' would be near-vacuous
Could not verify: Could not verify \_csharp\_method\_receiver\_types.
The verifier did not have enough to check \_csharp\_method\_receiver\_types, so it is saying so rather than guessing. No false assurance is the whole point.
Guarantee: No guarantee either way, this is an honest abstention, not a pass.
Note: Reason: no capturable inputs from the test suite; property tier: not verifiable: all 233 sampled inputs raised on both versions — the function never executed, so 'no divergence' would be vacuous (mostly AttributeError — names the real obstacle, not a sampling gap)
No difference found (not proven): No behavior difference found in \_csharp\_scoped\_receiver\_type (not a proof).
The verifier ran both versions of \_csharp\_scoped\_receiver\_type on many inputs and saw identical behavior every time. Strong evidence the change is safe, but evidence, not a proof.
Guarantee: Empirical: concolic exploration (CrossHair). A divergence on an untested input remains possible, so this is 'no counterexample found', not 'proven equivalent'.
Note: An input the sampler did not try could still differ.
Could not verify: Could not verify \_extract\_generic.
The verifier did not have enough to check \_extract\_generic, so it is saying so rather than guessing. No false assurance is the whole point.
Guarantee: No guarantee either way, this is an honest abstention, not a pass.
Note: Reason: no capturable inputs from the test suite; property tier: parameter `config` is annotated `LanguageConfig` — outside the synthesizable primitive/collection set
· 1 grounded finding(s) anchored inline below; 2 more finding(s) on lines outside this diff (see the check run).
Closes #4266
var report = ServiceFactory.Get<IReport>(scope); report.Build()got nocallsedge toIReport.Build, while the same code withIReport report = ...did. Avarwas only typed when its initializer wasnew T(), and the factory usually lives in another file, so the extractor can't see its return type anyway.Now the C# extractor records
var x = R.M<A>(...)(withRa type name, a typed local/field/parameter,this, or no receiver) as a deferred receiver[R, M, type args]on calls made throughx, and each method's declared return type and type-parameter count goes into the per-file tables C# already exports for the cross-file pass._resolve_csharp_member_callsfindsMonR(or its bases) the same way it finds any member call, and typesxasM's return type:T Get<T>()called asGet<IReport>()givesIReport,List<T> CreateList<T>()givesList, andWidget Make()givesWidget. There's no edge ifMisn't declared in the corpus, if more than one overload is left after matching the number of type arguments, if the type argument is inferred, or if the return type is a class-level type parameter (Box<T>.Value()).Limits: only a plain invocation as the initializer is handled.
await,?., chained calls (a.B().C()) and a generic receiver (Factory<T>.Get()) stay untyped. The variable's type is resolved with the same name lookup typed locals already use, so it carries over that lookup's existing gaps. On Polly, calls through a non-genericResiliencePipelineend up onResiliencePipeline<T>'s method of the same name, because the two types share one node (#4249); an explicitly typedResiliencePipeline pdoes the same on v8.Testing: new
tests/test_csharp_var_call_return_type.pywith 9 tests. The 4 positive ones (generic factory in the same file and across files, non-generic return,List<T>return) fail on current v8. The overload, missing-declaration, inferred-generic, class-type-parameter and explicit-type/newcontrol tests make sure no wrong edges show up. The full suite on 3.10/3.12/3.13/3.14 has the same failures as v8. Ruff is clean and pyright shows no new findings. The issue's repro now givesConsumer.Run() -> IReport.Build()for both the one-file and two-file layouts. On MediatR, Polly, Newtonsoft.Json and AutoMapper the C#callsedges went from 477 to 477, 5284 to 5483, 14419 to 14423 and 3526 to 3534, with none removed. I checked a sample of the new edges against the source.AI-assisted (Grok Bot); each commit carries a Co-Authored-By line.