Skip to content

feat: replace GPT-5.4-mini experiments with GPT-5.6 Luna - #269

Merged
Rodriguespn merged 1 commit into
mainfrom
Rodriguespn/ai-1190-replace-54-mini-with-56-luna
Sep 9, 2026
Merged

feat: replace GPT-5.4-mini experiments with GPT-5.6 Luna#269
Rodriguespn merged 1 commit into
mainfrom
Rodriguespn/ai-1190-replace-54-mini-with-56-luna

Conversation

@Rodriguespn

@Rodriguespn Rodriguespn commented Sep 7, 2026

Copy link
Copy Markdown
Contributor

What

Replaces the four GPT-5.4-mini experiments with GPT-5.6 Luna, the new Assistant default model, so the model change in supabase/supabase#49749 is covered by the public eval suite.

Why

The Assistant default model changed to gpt-5.6-luna (supabase/supabase#49749); the public evals suite should track that model.

Closes AI-1190

@vercel

vercel Bot commented Sep 7, 2026

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
evals Ready Ready Preview Sep 9, 2026 5:14pm UTC

Request Review

@Rodriguespn Rodriguespn added the run-evals Add to a PR to refresh benchmark evals label Sep 7, 2026
@Rodriguespn
Rodriguespn marked this pull request as ready for review September 7, 2026 17:35
@Rodriguespn
Rodriguespn requested a review from a team September 7, 2026 17:35
Comment thread experiments/codex-gpt-5.6-luna-no-skills.ts
@Rodriguespn Rodriguespn removed the run-evals Add to a PR to refresh benchmark evals label Sep 8, 2026
Rodriguespn added a commit that referenced this pull request Sep 8, 2026
Clarifies naming alongside the new codex-gpt-5.6-luna experiments, per
review feedback on #269. Pure rename: model, suite, skills, and runtime
config are unchanged.
Replaces the four GPT-5.4-mini experiments with GPT-5.6 Luna, the new
Assistant default model (supabase/supabase#49749), using
reasoningEffort medium to match. Renames the existing codex-gpt-5.6
experiments to codex-gpt-5.6-sol for naming clarity alongside the new
luna experiments (per PR review). Removes 8 experiments never wired
into any experiment suite (benchmark/no-skills/regression): they
carry no historical results and no live references.

Refreshed eval-results.json / regression-eval-results.json to match
via the run-evals CI label.
@Rodriguespn
Rodriguespn force-pushed the Rodriguespn/ai-1190-replace-54-mini-with-56-luna branch from a35be8c to 6cd27a7 Compare September 9, 2026 17:14
@Rodriguespn
Rodriguespn merged commit d462078 into main Sep 9, 2026
7 checks passed
mattrossman added a commit that referenced this pull request Sep 9, 2026
Moves the nightly regression suite from Claude Code / Sonnet 5 to Codex
/ GPT-5.6 Luna at medium effort, reusing the existing experiments from
#269 . See
[thread](https://supabase.slack.com/archives/C051L8U2EJF/p1788983215168069)
for context on the decision.

This PR doesn't update results, that will come from the next scheduled
run.

Closes AI-1200

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants