Skip to content

Repository files navigation

title ObviousBench v0.2.0 Public-Local Release Bundle
date 2026-08-13
type release-bundle
status local-prep

ObviousBench v0.2.0 Public-Local Release Bundle

This bundle is allowlisted for public review but has not been published. It contains public examples plus the aggregate v0.2 summary CSV and report.

It intentionally excludes private held-out prompts, raw logs, raw model outputs, private review HTML, private item-level outcomes, and private attempt-level outcomes.

Snapshot

  • Private held-out items: 144
  • Model/config rows: 565
  • Complete rows: 565
  • Attempt rows: 244080
  • Scored attempts: 244080
  • Estimated cost: $259.07

Included Files

  • configs/registries/model_registry_v1.yaml
  • configs/registries/model_thinking_settings_v1.yaml
  • data/public_examples/arithmetic.jsonl
  • data/public_examples/character_count.jsonl
  • data/public_examples/constraint_awareness.jsonl
  • data/public_examples/format_compliance.jsonl
  • data/public_examples/negation.jsonl
  • data/public_examples/ordering.jsonl
  • data/public_examples/spelling_transform.jsonl
  • data/public_examples/word_count.jsonl
  • docs/positioning/background-and-rhetoric.md
  • docs/reference/benchmark-card.md
  • docs/reference/methodology.md
  • docs/reference/scoring-policy.md
  • docs/reference/source-policy.md
  • docs/release/v0_2/generated/README.md
  • docs/release/v0_2/generated/background-and-rhetoric.md
  • docs/release/v0_2/generated/github-release-notes.md
  • docs/release/v0_2/generated/huggingface-dataset-card.md
  • docs/release/v0_2/generated/interactive-results.html
  • docs/release/v0_2/generated/launch-essay-draft.md
  • docs/release/v0_2/generated/model-size-review.csv
  • docs/release/v0_2/generated/project-page.md
  • docs/release/v0_2/generated/provenance.json
  • docs/release/v0_2/generated/public-release-checklist.md
  • docs/release/v0_2/generated/release-metadata.json
  • docs/release/v0_2/generated/social-snippets.md
  • reports/v0_2/aggregate/report.md
  • reports/v0_2/aggregate/summary.csv

About

A lightweight Inspect AI benchmark for obvious public-facing LLM failures.

Topics

Resources

Stars

4 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages