Skip to content

external_results: ValorBrain BEAM-100K (80.8% / 70.9%) - #36

Open
agentxagi wants to merge 2 commits into
vectorize-io:mainfrom
agentxagi:feat/valorbrain-beam-100k-results
Open

external_results: ValorBrain BEAM-100K (80.8% / 70.9%)#36
agentxagi wants to merge 2 commits into
vectorize-io:mainfrom
agentxagi:feat/valorbrain-beam-100k-results

Conversation

@agentxagi

Copy link
Copy Markdown

What

Adds two ValorBrain runs to the beam/100k section of external_results.json:

Memory Accuracy Reader Judge
ValorBrain 80.8% (323/400) stealth/ox-alpha @ reasoning.effort=max glm-5.2
ValorBrain (GLM-5.2 reader) 70.9% (312/400) glm-5.2 Gemini 3.6 Flash

Reproducibility

Same model as Hindsight/mem0-cloud: proprietary engine, public API, reproducible results.

Notes

  • Both runs are the full 400-query beam/100k split (20 conversations).
  • The 80.8% run's reader is stealth/ox-alpha via OpenRouter with reasoning.effort: max; retrieval is dense halfvec HNSW + BM25 hybrid inside PostgreSQL.
  • Per-category breakdown for both runs is in the writeups and the provider repo README.

Happy to adjust format/labels to match how you ingest external entries.

Two runs on the beam/100k split (400 queries, 20 conversations):

- ValorBrain: 80.8% (323/400) — stealth/ox-alpha reader at
  reasoning.effort=max, glm-5.2 judge. First run above 80% on this
  harness to our knowledge.
- ValorBrain (GLM-5.2 reader): 70.9% (312/400) — glm-5.2 reader,
  Gemini 3.6 Flash judge.

Both reproducible via the ValorBrain provider in ValorBrain/valorbrain-amb
(public API, same model as Hindsight/mem0-cloud). Full writeups linked
in each entry's source_url.
@vercel

vercel Bot commented Aug 24, 2026

Copy link
Copy Markdown

Someone is attempting to deploy a commit to the Vectorize Team on Vercel.

A member of the Team first needs to authorize it.

200 queries, 10 giant documents (~110M tokens total). stealth/ox-alpha
reader at reasoning.effort=max, glm-5.2 judge — same config as our
BEAM-100K 80.8% run. Runner accuracy 51.2% (112/200 binary correct),
+10.6pp over Hindsight RAG (40.6%) on this split. Full run artifact:
results/beam-10m-ox-alpha-effortmax-200q.json in ValorBrain/valorbrain-amb.
@agentxagi

Copy link
Copy Markdown
Author

Model identity correction: the stealth/ox-alpha reader used in these runs is GLM-5.3-flash (Z.ai), served via OpenRouter under the stealth alias. The scores stand as measured; attribution updated here for the record.

Notable: the reader and judge are both Z.ai models (GLM-5.3-flash reader, glm-5.2 judge). The judge scores against gold answers with rubrics, and the same configuration was used across all runs (including the 70.9% GLM-5.2-reader baseline), so comparisons remain internally consistent. A flash-tier open-weights model beating Gemini-based readers at this price class is the headline.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant