Skip to content

feat(recall): cap tag-score denominator to fix query-length bias - #193

Merged
jack-arturo merged 3 commits into
developfrom
feat/tag-score-token-cap
Jun 11, 2026
Merged

jack-arturo merged 3 commits into
developfrom
feat/tag-score-token-cap

Conversation

@jack-arturo

Copy link
Copy Markdown
Member

Continuation of #185, which GitHub auto-closed (and refused to reopen) when its stacked base branch feat/tunable-recency was deleted on #182's merge. Identical branch and content, now rebased onto develop with #182 included.

See #185 for the full description, validation evidence (production-corpus lab A/B that flipped the default to SEARCH_TAG_SCORE_TOKEN_CAP=0), and review history.

Part of the develop-branch integration series for the ranking release.

🤖 Generated with Claude Code

jack-arturo and others added 3 commits June 11, 2026 20:53
The tag-overlap score divided token hits by the full query length, so a
1-token query earned full tag credit from a single hit while a 5-token
query needed all 5 tokens to match — biasing short queries up and long
queries down.

Introduce SEARCH_TAG_SCORE_TOKEN_CAP (default 3): the denominator
becomes min(len(query_tokens), cap), clipped to 1.0 on the way out.
Setting the cap to 0 restores the legacy full-query-length denominator;
negative values fall back to the default (safer than treating a typo'd
negative as an intentional legacy opt-out), and unparseable values raise
like the neighboring int()/float() env parses.

This intentionally changes default scoring for queries longer than 3
tokens. The default-change ships behind the section-4 eval gates per the
release plan; the legacy escape hatch (cap=0) is covered by tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… production-corpus regression evidence

Production-corpus A/B testing (2026-06-11, 200 queries against a
10,107-memory production clone, 3-run baseline R@5 0.650-0.655 /
MRR 0.429-0.433) showed every tested cap value regresses recall on
ungated free-text queries:

- cap=2: R@5 0.510 (-14.2pp, paired p<0.0001), MRR 0.324
- cap=3 (previous default): R@5 0.580 (-7.2pp, p=0.0002), MRR 0.361
- cap=4: R@5 0.615 (-3.7pp, p=0.0103), MRR 0.390
- cap=0 (legacy) = baseline

Mechanism: on ungated free-text queries the capped denominator inflates
tag scores (1 matching tag on a 12-token query: 0.083 -> 0.33),
amplifying tag noise over vector/keyword evidence. A two-stack probe
A/B also showed cap=3 raising top-1 scores on known-garbage negative
probes by +0.04.

The cap remains available as an opt-in for tag-scoped retrieval
experiments; all cap-behavior tests already pin the value explicitly
via monkeypatch, so no test expectations changed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Copilot AI review requested due to automatic review settings June 11, 2026 19:02
@jack-arturo
jack-arturo merged commit cefa516 into develop Jun 11, 2026
3 checks passed
@jack-arturo
jack-arturo deleted the feat/tag-score-token-cap branch June 11, 2026 19:02

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR introduces an optional query-length normalization for the tag/metadata overlap component in recall scoring by capping the tag-score denominator via a new env var (SEARCH_TAG_SCORE_TOKEN_CAP). It’s designed to mitigate query-length bias while preserving legacy behavior by default (0).

Changes:

  • Add SEARCH_TAG_SCORE_TOKEN_CAP config and parsing helper (_non_negative_int_or_default) with default 0 (legacy behavior).
  • Update _compute_metadata_score() to use a capped denominator when enabled and clamp tag score to <= 1.0.
  • Add targeted unit tests covering cap behavior, legacy mode, clipping, and config parsing; document the new env var.

Reviewed changes

Copilot reviewed 4 out of 4 changed files in this pull request and generated 1 comment.

File Description
automem/utils/scoring.py Applies optional capped denominator for tag score and clamps to 1.0 to preserve 0–1 component bounds.
automem/config.py Adds _non_negative_int_or_default() and introduces SEARCH_TAG_SCORE_TOKEN_CAP env configuration.
tests/test_api_endpoints.py Adds/updates tests to pin cap behavior and validate cap/legacy/clipping and config-domain parsing.
docs/ENVIRONMENT_VARIABLES.md Documents the new SEARCH_TAG_SCORE_TOKEN_CAP environment variable.

| `SEARCH_WEIGHT_METADATA` | Metadata sidecar match | `0.35` | Candidates admitted via the metadata sidecar channel (see `RECALL_METADATA_SEARCH_ENABLED`) |
| `SEARCH_WEIGHT_RELATION` | Graph relationship boost | `0.25` | Memories connected via edges |
| `SEARCH_WEIGHT_TAG` | Tag matching | `0.20` | Tag overlap scoring |
| `SEARCH_TAG_SCORE_TOKEN_CAP` | Tag-score denominator cap | `0` (opt-in) | When > 0, tag score divides token hits by `min(query tokens, cap)`; `0` (default) keeps the legacy full-query-length denominator. Production-corpus A/B (2026-06-11, 200 queries, 10k-memory clone) showed cap values 2/3/4 regress Recall@5 by 14/7/4pp on ungated queries; enable only for tag-scoped retrieval experiments. Not a weight |
jack-arturo added a commit that referenced this pull request Jun 12, 2026
…nce gate, date-aware ranking (#182, #193, #186, #187, #183, #184, #188) (#194)

## Release: ranking & recall series (develop → main)

⚠️ **Merge with a MERGE COMMIT — do not squash.** release-please needs
the individual conventional commits below to compute the version and
changelog for PR #154.

### What's in this release

| PR | Change | Default behavior |
|---|---|---|
| #182 | `feat(recall)`: configurable recency decay window/curve |
unchanged (env-gated) |
| #193 (replaces #185) | `feat(recall)`: tag-score denominator cap fixes
query-length bias | unchanged (`SEARCH_TAG_SCORE_TOKEN_CAP=0`) |
| #186 | `fix(recall)`: relevance gate — query-independent scoring gated
on topical evidence (#130) | unchanged (gate off) |
| #187 | `feat(recall)`: date-aware ranking,
`recency_bias=off\|on\|auto`, latest-fact selection (#158, #159) |
`RECALL_RECENCY_BIAS=off`; adds deterministic timestamp tiebreak for
near-ties |
| #183 | `feat(benchmarks)`: failure-mode diagnosis harness + judge
quota preflight | tooling only |
| #184 | `fix(mcp)`: surface stored metadata + `updated_at` in detailed
recall format (#111) | additive |
| #188 | `feat(enrichment)`: classification fallback-rate metrics in
`/enrichment/status` | additive |

Plus: CI now runs on `develop` pushes/PRs; benchmark experiment log +
README contribution-policy note.

### Verification evidence

- **Unit/lint/npm**: 625 pytest + 16 mcp-sse-server tests green on
develop head; CI green.
- **Default-preserve**: recall-lab baseline on the 10k-memory production
snapshot — develop defaults vs main pooled baseline identical aggregates
(R@5 0.655 / R@10 0.710 / MRR 0.434 / NDCG@10 0.501). Two-stack probe
run (main vs develop, defaults): 11/12 preserve-exact, remaining diffs
are near-tie reorders (top-1 score deltas ≤ 5.4e-5, the #187 timestamp
tiebreak).
- **Full judged 500q LongMemEval** (ship config:
`RECALL_RECENCY_BIAS=auto` + `temporal-answer` harness): recall@5 96.6%
(483/500), accuracy 86.0% (430/500), `judge_errors=0`,
`memory_ingest_failures=0`.
- **Churn attribution** (targeted re-runs of all 17 churned questions on
current-main-at-defaults and develop-at-defaults): 15/17 moved with #191
(already on main) — the April canonical 97.2% floor is stale; current
main measures ~97.0%. Develop-at-defaults differs from current main by
**1 question in 500** (a near-tie rank-5/6 flip from #187's
deterministic tiebreak). Accuracy is within answerer replicate noise
(identical-config reference runs flip 28/500 answers).
- Full detail: `benchmarks/EXPERIMENT_LOG.md` (2026-06-11 entry) and
`benchmarks/results/lme_churn17_*` + `analyze_churn17.py`.

### Opt-in features shipped OFF

`RECALL_RELEVANCE_GATE` (validated at 0.40 on lab corpus; improves
negative-probe precision) and `RECALL_RECENCY_BIAS=auto` (current-state
query re-ranking). Neither affects default behavior; see
`docs/ENVIRONMENT_VARIABLES.md`.

### After merging

release-please will update PR #154 (v0.16.0); merging *that* cuts the
tag and publishes the `:stable` image — the actual user-facing deploy
event for Railway template users.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
jack-arturo added a commit that referenced this pull request Jun 26, 2026
🤖 I have created a release *beep* *boop*
---


##
[0.16.0](v0.15.2...v0.16.0)
(2026-06-26)


### Features

* **api:** add admin backup endpoint
([#162](#162))
([8b1f264](8b1f264))
* **api:** support bulk memory associations
([1221e36](1221e36))
* **api:** support bulk memory associations
([#198](#198))
([28eb916](28eb916))
* **benchmarks:** LongMemEval failure-mode diagnosis harness + judge
quota preflight
([#183](#183))
([f99bece](f99bece))
* **consolidation:** expose cluster threshold and min size as env vars
([#163](#163))
([7e731f3](7e731f3))
* **enrichment:** expose classification fallback-rate metrics in
/enrichment/status
([#188](#188))
([0b522a9](0b522a9))
* **entity:** harden identity cleanup and repair tooling
([#176](#176))
([827dfbc](827dfbc))
* **eval:** recall-quality optimization harness — lab foundation +
design ([#197](#197))
([431433e](431433e))
* **graph:** support unbounded visualizer snapshots
([#141](#141))
([c730128](c730128))
* **lab:** add aged labelled distractor injection
([cc5d546](cc5d546))
* **lab:** add config_complexity simplicity metric
([dfb10d9](dfb10d9))
* **lab:** add distractor_rate_at_k precision guardrail metric
([872eab2](872eab2))
* **lab:** add lab_corpus with parameterized recall
([5e1e071](5e1e071))
* **lab:** add pick_winner scorecard decision rule
([3187eac](3187eac))
* **lab:** add real consolidation pass helper
([48a7d4a](48a7d4a))
* **lab:** isolate production clone restores
([#171](#171))
([aef90c0](aef90c0))
* **lab:** wire scorecard, distractors, recall params, consolidation
into runner
([589ec30](589ec30))
* **recall:** add metadata sidecar search
([#177](#177))
([4e7956e](4e7956e))
* **recall:** add state_mode=current|history recall alias
([#173](#173))
([b1df86c](b1df86c))
* **recall:** cap tag-score denominator to fix query-length bias
([#193](#193))
([cefa516](cefa516))
* **recall:** date-aware ranking + latest-fact selection
([#158](#158),
[#159](#159))
([#187](#187))
([a6ed945](a6ed945))
* **recall:** make recency decay window and curve configurable
([#182](#182))
([dbb933f](dbb933f))
* **recall:** ranking release — recency config, tag-score cap, relevance
gate, date-aware ranking
([#182](#182),
[#193](#193),
[#186](#186),
[#187](#187),
[#183](#183),
[#184](#184),
[#188](#188))
([#194](#194))
([337fe98](337fe98))
* **scripts:** safer reclassify_with_llm.py with provider flags +
tighter prompt
([#164](#164))
([a742602](a742602))


### Bug Fixes

* **api:** address copilot review on PR
[#198](#198)
([0466a1e](0466a1e))
* **api:** handle grouped association write failures
([cd93df9](cd93df9))
* **backup:** make backup_automem.py runnable as `python
scripts/backup_automem.py`
([#175](#175))
([edd9742](edd9742))
* **benchmarks:** add publication verification bundle
([#166](#166))
([420d721](420d721))
* **consolidation:** skip eager first tick at startup to avoid FalkorDB
load race
([#165](#165))
([1b812cf](1b812cf))
* **docs:** keep dispatch payload arrays stable
([df6e9e8](df6e9e8))
* **embedding:** fall back to per-item real embeddings before
placeholders in batch path
([#189](#189))
([6e9c62c](6e9c62c))
* **entity:** restore person-shape exemption on the slug validation path
([#179](#179))
([5e29960](5e29960))
* **entity:** stop validator over-rejecting real people, code tools, and
event categories
([#178](#178))
([193b730](193b730))
* **lab:** address copilot review on PR
[#197](#197)
([45f80d6](45f80d6))
* **lab:** align scorecard key contract (build_scorecard -&gt;
pick_winner)
([7d91530](7d91530))
* **mcp-sse:** decouple /health liveness from upstream readiness
([#151](#151))
([5bcfb8b](5bcfb8b))
* **mcp:** cap association failure summary
([ea4e08f](ea4e08f))
* **mcp:** surface stored metadata and updated_at in detailed recall
format ([#184](#184))
([230416e](230416e))
* **recall:** address copilot review on PR
[#194](#194)
([50b1647](50b1647))
* **recall:** canonicalize / and : separators in context_tag matching
([3afd9d3](3afd9d3))
* **recall:** canonicalize / and : separators in context_tag matching
([#203](#203))
([ba5e9ff](ba5e9ff))
* **recall:** gate query-independent scoring on topical evidence within
tag scope
([#130](#130))
([#186](#186))
([c11b594](c11b594))
* **recall:** hydrate semantic recall summaries
([#192](#192))
([76e845d](76e845d))
* **recall:** normalize graph keyword scores into the 0-1 component
range ([#191](#191))
([3653ddf](3653ddf))
* **recall:** respect current memory state
([#170](#170))
([ed36b98](ed36b98)),
closes [#169](#169)
[#158](#158)
[#159](#159)
* **scripts:** add sys.path guard to reembed_embeddings.py
([d333cf0](d333cf0))


### Documentation

* add scripts catalog, recall-quality-lab guide, and 0.16.0 migrations
([f20c664](f20c664))
* **bench:** log full judged 500q LongMemEval ship-config run with churn
attribution
([41bf8d0](41bf8d0))
* **eval:** Plan A — lab metric foundation (TDD, 9 tasks)
([0087dda](0087dda))
* **eval:** Plan B — parallel matrix harness (TDD, 9 tasks)
([c8ddfb2](c8ddfb2))
* **evals:** mark Memora/FAMA/WRIT lifecycle diagnostics as
diagnostic-only
([#174](#174))
([e8a3285](e8a3285))
* **eval:** spec for recall-quality optimization harness
([b1a1995](b1a1995))
* fix stale claims and document gated flags for 0.16.0
([b152d64](b152d64))
* note develop-branch contribution policy in README
([ccf02dd](ccf02dd))
* **positioning:** add scout reference
([#168](#168))
([922d23b](922d23b))
* refresh benchmark currency for the neutral AMB run and prune stale
archive docs
([3ff95bd](3ff95bd))
* refresh benchmark currency for the neutral AMB run and prune stale
archive docs
([#204](#204))
([89c30e0](89c30e0))
* refresh README and benchmark guidance
([#157](#157))
([bba31cc](bba31cc))
* **runtime:** align Docker viewer paths and setup guidance
([#155](#155))
([bbda79b](bbda79b))
* scripts catalog, recall quality lab guide, and 0.16.0 migration
runbook ([#199](#199))
([f190ae5](f190ae5))

---
This PR was generated with [Release
Please](https://github.com/googleapis/release-please). See
[documentation](https://github.com/googleapis/release-please#release-please).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants