Repository navigation
Derive data-flow node locations & sink taints from storage (deterministic taint findings) - #11952
Conversation
Taint resolution previously ran a fixed 40-round BFS and stopped at a hard-coded nesting limit, so flows longer than 40 hops were silently missed and the round count was unrelated to convergence. Run to a true fixed point instead: loop while sources and sinks remain, relying on the existing (id, taints) visited guard -- whose state space is finite -- to guarantee termination. This removes the resolution-depth limit entirely, so deep taint flows are no longer missed. Propagation semantics are otherwise unchanged: each (id, taints) state is explored exactly as before. Adds tests for a flow deeper than the old 40-hop cap (via plain assignments and via a chain of distinct specialized function calls) and for a sanitized array value that must not be reported. Co-Authored-By: Claude Opus 4.8
Restrict resolution to the sub-graph from which a sink is actually reachable, via a backward reachability pass from the sinks over the forward edges: a node from which no sink is reachable can never produce an issue, so propagating taint into it is wasted work. On real codebases the full taint graph is huge while this relevant sub-graph is tiny, which is what keeps running to convergence cheap. Specialized and unspecialized nodes are linked both by their authoritative unspecialized_id field and by the id-string separator, so the prune never drops a node whose specialized or de-specialized form can reach a sink (the resolution walk maps freely between the two). Co-Authored-By: Claude Opus 4.8
When analysis runs across forked workers, each builds a partial DataFlowGraph the parent merges with `+=` (first value per key wins). The same node id could be created at several sites with different code_locations and sink taints, so which survived the merge depended on worker scheduling. Since a taint issue is reported at its node's code_location, this shifted reported locations and per-file occurrence counts between identical runs, so baselined issues silently reappeared and --set-baseline never converged. Make a node's location and sink taints a pure function of its identity: * getForCallableArg()/getForCallableReturn() no longer accept an independent CodeLocation (breaking). A callable node has no storage to derive a canonical location from, so its location is now always its specialization (the callsite), which is already part of the node id. * getForMethodArgument() derives its sink taints from the parameter's storage rather than a caller-supplied argument (breaking). * The inherited-method return-node kind is dropped: MethodCallReturnType- Fetcher always builds return nodes via getForMethodReturn(), resolving the declaring storage, so the location is the canonical return location wherever the id is produced. Co-Authored-By: Claude Opus 4.8
The argument sink node id (`Class::method#offset`) is shared by every call of the method, so its location must be a function of that identity. When the caller has no storage in hand (e.g. an inherited method has no storage under its own id), resolve the declaring method's storage from the cased method id itself, rather than leaving the location unset here and letting it differ between the forked-worker graphs. Every site that mints the node then agrees on the canonical parameter location. Co-Authored-By: Claude Opus 4.8
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit 4a6385c. Configure here.
| : ($method_id instanceof MethodIdentifier ? 'magic-method' : 'callable-object'), | ||
| $cased_method_id, | ||
| $argument_offset, | ||
| $function_param->location ?? $code_location, |
There was a problem hiding this comment.
Sink nodes skip storage lookup
Medium Severity
In the same checkArgumentsMatch pass, ArgumentAnalyzer::processTaintedness can resolve declaring method storage and build argument nodes with getForMethodArgument, but the taint-sink block still uses getForCallableArg whenever $function_storage is null. Shared node ids then get canonical parameter metadata on flow edges and callsite-oriented metadata on sinks, so forked graph merges can remain nondeterministic for annotated sinks.
Reviewed by Cursor Bugbot for commit 4a6385c. Configure here.


When analysis runs across forked workers, each builds a partial
DataFlowGraphthe parent merges with+=(first value per key wins). The sameDataFlowNodeid could be created at several sites with differentcode_locations and sink taints, so which survived the merge depended on worker scheduling. Since a taint issue is reported at its node'scode_location, this shifted reported locations and per-file occurrence counts between otherwise identical runs — baselined issues silently reappeared and--set-baselinenever converged.This makes a node's location and sink taints a pure function of its identity:
getForCallableArg()/getForCallableReturn()no longer accept an independentCodeLocation(breaking). A callable node has no storage to derive a canonical location from, so its location is now always its specialization (the callsite), which is already part of the node id.getForMethodArgument()derives its sink taints from the parameter's storage rather than a caller-supplied argument (breaking).inherited-methodreturn-node kind is dropped:MethodCallReturnTypeFetcheralways builds return nodes viagetForMethodReturn(), resolving the declaring storage.Verified on a ~9k-file project:
--set-baselinefollowed by repeated analysis now reports zero new issues across many consecutive runs. TaintTest passes unchanged.Part 2 of a 4-PR stack. Stacked on #11951 — review/merge that first; the diff here includes its commits until it merges.
🤖 Generated with Claude Code