Skip to content

feat(eval): filter reported scores and metrics - #130

Merged
Abhijeet Prasad (AbhiPrasad) merged 3 commits into
mainfrom
abhi-feat-sdk-196-filter-pr-metrics
Aug 20, 2026
Merged

feat(eval): filter reported scores and metrics#130
Abhijeet Prasad (AbhiPrasad) merged 3 commits into
mainfrom
abhi-feat-sdk-196-filter-pr-metrics

Conversation

@AbhiPrasad

Copy link
Copy Markdown
Member

Add report_scores and report_metrics action inputs to constraint PR comment output for eval action.

Example:

- uses: braintrustdata/eval-action@v2
  with:
    api_key: ${{ secrets.BRAINTRUST_API_KEY }}
    runtime: node
    report_scores: Levenshtein, Factuality
    report_metrics: |
      Duration
      Cost

Resolves https://linear.app/braintrustdata/issue/SDK-196/add-a-method-to-specify-which-metrics-are-reported-in-the-pr-comment

Add independent `report_scores` and `report_metrics` action inputs so PR
comments can focus on selected evaluation results. Render scores and metrics in
separate tables while preserving the existing report-all behavior when filters
are omitted.

Resolves https://linear.app/braintrustdata/issue/SDK-196/add-a-method-to-specify-which-metrics-are-reported-in-the-pr-comment
Trigger eval workflows from the pull request event so the action has a PR to
comment on. Limit push-triggered evals to `main` to avoid duplicate runs on
feature branches.
@github-actions

github-actions Bot commented Aug 19, 2026

Copy link
Copy Markdown

Braintrust eval report

Say Hi Bot Python (abhijeet@braintrustdata.com-1787181833)

Name Average Improvements Regressions
Scores
Levenshtein 77.8% (+0pp) - -
Metrics
Llm_calls 0 (+0) - -
Tool_calls 0 (+0) - -
Errors 0 (+0) - -
Llm_errors 0 (+0) - -
Tool_errors 0 (+0) - -
Prompt_tokens 0tok (+0tok) - -
Prompt_cached_tokens 0tok (+0tok) - -
Prompt_cache_creation_tokens 0tok (+0tok) - -
Prompt_cache_creation_5m_tokens 0tok (+0tok) - -
Prompt_cache_creation_1h_tokens 0tok (+0tok) - -
Completion_tokens 0tok (+0tok) - -
Completion_reasoning_tokens 0tok (+0tok) - -
Total_tokens 0tok (+0tok) - -
Duration 0s (0s) 2 🟢 -

@github-actions

github-actions Bot commented Aug 19, 2026

Copy link
Copy Markdown

Braintrust eval report

Say Hi Bot Python (abhijeet@braintrustdata.com-1787181835)

Name Average Improvements Regressions
Scores
Levenshtein 77.8% (+0pp) - -
Metrics
Llm_calls 0 (+0) - -
Tool_calls 0 (+0) - -
Errors 0 (+0) - -
Llm_errors 0 (+0) - -
Tool_errors 0 (+0) - -
Prompt_tokens 0tok (+0tok) - -
Prompt_cached_tokens 0tok (+0tok) - -
Prompt_cache_creation_tokens 0tok (+0tok) - -
Prompt_cache_creation_5m_tokens 0tok (+0tok) - -
Prompt_cache_creation_1h_tokens 0tok (+0tok) - -
Completion_tokens 0tok (+0tok) - -
Completion_reasoning_tokens 0tok (+0tok) - -
Total_tokens 0tok (+0tok) - -
Duration 0s (0s) 2 🟢 -

@github-actions

github-actions Bot commented Aug 19, 2026

Copy link
Copy Markdown

Braintrust eval report

Say Hi Bot (HEAD-1787181841)

Name Average Improvements Regressions
Scores
Levenshtein 100% (+0pp) - -
Metrics
Llm_calls 0 (+0) - -
Tool_calls 0 (+0) - -
Errors 0 (+0) - -
Llm_errors 0 (+0) - -
Tool_errors 0 (+0) - -
Prompt_tokens 0tok (+0tok) - -
Prompt_cached_tokens 0tok (+0tok) - -
Prompt_cache_creation_tokens 0tok (+0tok) - -
Prompt_cache_creation_5m_tokens 0tok (+0tok) - -
Prompt_cache_creation_1h_tokens 0tok (+0tok) - -
Completion_tokens 0tok (+0tok) - -
Completion_reasoning_tokens 0tok (+0tok) - -
Total_tokens 0tok (+0tok) - -
Duration 1s (+0s) 3 🟢 5 🔴

@github-actions

github-actions Bot commented Aug 19, 2026

Copy link
Copy Markdown

Braintrust eval report

Console logging (HEAD-1787181836)

Name Average Improvements Regressions
Scores
Levenshtein 100% (+0pp) - -
Metrics
Llm_calls 0 (+0) - -
Tool_calls 0 (+0) - -
Errors 0 (+0) - -
Llm_errors 0 (+0) - -
Tool_errors 0 (+0) - -
Prompt_tokens 0tok (+0tok) - -
Prompt_cached_tokens 0tok (+0tok) - -
Prompt_cache_creation_tokens 0tok (+0tok) - -
Prompt_cache_creation_5m_tokens 0tok (+0tok) - -
Prompt_cache_creation_1h_tokens 0tok (+0tok) - -
Completion_tokens 0tok (+0tok) - -
Completion_reasoning_tokens 0tok (+0tok) - -
Total_tokens 0tok (+0tok) - -
Duration 0s (0s) 5 🟢 2 🔴

My Evaluation (HEAD-1787181836)

Name Average Improvements Regressions
Scores
Exact match 100% (+0pp) - -
Metrics
Llm_calls 0 (+0) - -
Tool_calls 0 (+0) - -
Errors 0 (+0) - -
Llm_errors 0 (+0) - -
Tool_errors 0 (+0) - -
Prompt_tokens 10tok (+0tok) - -
Prompt_cached_tokens 0tok (+0tok) - -
Prompt_cache_creation_tokens 0tok (+0tok) - -
Prompt_cache_creation_5m_tokens 0tok (+0tok) - -
Prompt_cache_creation_1h_tokens 0tok (+0tok) - -
Completion_tokens 2tok (+0tok) - -
Completion_reasoning_tokens 0tok (+0tok) - -
Total_tokens 12tok (+0tok) - -
Duration 1s (+0.65s) - 1 🔴

Say Hi Bot (HEAD-1787181836)

Name Average Improvements Regressions
Scores
Levenshtein 100% (+0pp) - -
Metrics
Llm_calls 0 (+0) - -
Tool_calls 0 (+0) - -
Errors 0 (+0) - -
Llm_errors 0 (+0) - -
Tool_errors 0 (+0) - -
Prompt_tokens 0tok (+0tok) - -
Prompt_cached_tokens 0tok (+0tok) - -
Prompt_cache_creation_tokens 0tok (+0tok) - -
Prompt_cache_creation_5m_tokens 0tok (+0tok) - -
Prompt_cache_creation_1h_tokens 0tok (+0tok) - -
Completion_tokens 0tok (+0tok) - -
Completion_reasoning_tokens 0tok (+0tok) - -
Total_tokens 0tok (+0tok) - -
Duration 1s (0s) 4 🟢 5 🔴

@github-actions

github-actions Bot commented Aug 19, 2026

Copy link
Copy Markdown

Braintrust eval report

Say Hi Bot (HEAD-1787181849)

Name Average Improvements Regressions
Scores
Levenshtein 100% (+0pp) - -
Metrics
Llm_calls 0 (+0) - -
Tool_calls 0 (+0) - -
Errors 0 (+0) - -
Llm_errors 0 (+0) - -
Tool_errors 0 (+0) - -
Prompt_tokens 0tok (+0tok) - -
Prompt_cached_tokens 0tok (+0tok) - -
Prompt_cache_creation_tokens 0tok (+0tok) - -
Prompt_cache_creation_5m_tokens 0tok (+0tok) - -
Prompt_cache_creation_1h_tokens 0tok (+0tok) - -
Completion_tokens 0tok (+0tok) - -
Completion_reasoning_tokens 0tok (+0tok) - -
Total_tokens 0tok (+0tok) - -
Duration 1s (+0s) - 6 🔴

Keep scores and metrics visually distinct with labeled sections while sharing a
single table header for a more compact PR comment.
@AbhiPrasad
Abhijeet Prasad (AbhiPrasad) merged commit ed007eb into main Aug 20, 2026
9 checks passed
@github-actions

github-actions Bot commented Aug 20, 2026

Copy link
Copy Markdown

Braintrust eval report

Say Hi Bot Python (main-1787230576)

Name Average Improvements Regressions
Scores
Levenshtein 77.8% (+0pp) - -
Metrics
Llm_calls 0 (+0) - -
Tool_calls 0 (+0) - -
Errors 0 (+0) - -
Llm_errors 0 (+0) - -
Tool_errors 0 (+0) - -
Prompt_tokens 0tok (+0tok) - -
Prompt_cached_tokens 0tok (+0tok) - -
Prompt_cache_creation_tokens 0tok (+0tok) - -
Prompt_cache_creation_5m_tokens 0tok (+0tok) - -
Prompt_cache_creation_1h_tokens 0tok (+0tok) - -
Completion_tokens 0tok (+0tok) - -
Completion_reasoning_tokens 0tok (+0tok) - -
Total_tokens 0tok (+0tok) - -
Duration 0s (+0s) - 1 🔴

@github-actions

github-actions Bot commented Aug 20, 2026

Copy link
Copy Markdown

Braintrust eval report

Say Hi Bot Python (main-1787230581)

Name Average Improvements Regressions
Scores
Levenshtein 77.8% (+0pp) - -
Metrics
Llm_calls 0 (+0) - -
Tool_calls 0 (+0) - -
Errors 0 (+0) - -
Llm_errors 0 (+0) - -
Tool_errors 0 (+0) - -
Prompt_tokens 0tok (+0tok) - -
Prompt_cached_tokens 0tok (+0tok) - -
Prompt_cache_creation_tokens 0tok (+0tok) - -
Prompt_cache_creation_5m_tokens 0tok (+0tok) - -
Prompt_cache_creation_1h_tokens 0tok (+0tok) - -
Completion_tokens 0tok (+0tok) - -
Completion_reasoning_tokens 0tok (+0tok) - -
Total_tokens 0tok (+0tok) - -
Duration 0s (0s) 2 🟢 -

@github-actions

github-actions Bot commented Aug 20, 2026

Copy link
Copy Markdown

Braintrust eval report

Say Hi Bot (main-1787230582)

Name Average Improvements Regressions
Scores
Levenshtein 100% (+0pp) - -
Metrics
Llm_calls 0 (+0) - -
Tool_calls 0 (+0) - -
Errors 0 (+0) - -
Llm_errors 0 (+0) - -
Tool_errors 0 (+0) - -
Prompt_tokens 0tok (+0tok) - -
Prompt_cached_tokens 0tok (+0tok) - -
Prompt_cache_creation_tokens 0tok (+0tok) - -
Prompt_cache_creation_5m_tokens 0tok (+0tok) - -
Prompt_cache_creation_1h_tokens 0tok (+0tok) - -
Completion_tokens 0tok (+0tok) - -
Completion_reasoning_tokens 0tok (+0tok) - -
Total_tokens 0tok (+0tok) - -
Duration 1s (0s) 19 🟢 1 🔴

@github-actions

github-actions Bot commented Aug 20, 2026

Copy link
Copy Markdown

Braintrust eval report

Console logging (main-1787230587)

Name Average Improvements Regressions
Scores
Levenshtein 100% (+0pp) - -
Metrics
Llm_calls 0 (+0) - -
Tool_calls 0 (+0) - -
Errors 0 (+0) - -
Llm_errors 0 (+0) - -
Tool_errors 0 (+0) - -
Prompt_tokens 0tok (+0tok) - -
Prompt_cached_tokens 0tok (+0tok) - -
Prompt_cache_creation_tokens 0tok (+0tok) - -
Prompt_cache_creation_5m_tokens 0tok (+0tok) - -
Prompt_cache_creation_1h_tokens 0tok (+0tok) - -
Completion_tokens 0tok (+0tok) - -
Completion_reasoning_tokens 0tok (+0tok) - -
Total_tokens 0tok (+0tok) - -
Duration 0s (0s) 4 🟢 1 🔴

My Evaluation (main-1787230587)

Name Average Improvements Regressions
Scores
Exact match 100% (+0pp) - -
Metrics
Llm_calls 0 (+0) - -
Tool_calls 0 (+0) - -
Errors 0 (+0) - -
Llm_errors 0 (+0) - -
Tool_errors 0 (+0) - -
Prompt_tokens 10tok (+0tok) - -
Prompt_cached_tokens 0tok (+0tok) - -
Prompt_cache_creation_tokens 0tok (+0tok) - -
Prompt_cache_creation_5m_tokens 0tok (+0tok) - -
Prompt_cache_creation_1h_tokens 0tok (+0tok) - -
Completion_tokens 2tok (+0tok) - -
Completion_reasoning_tokens 0tok (+0tok) - -
Total_tokens 12tok (+0tok) - -
Duration 0.17s (-0.83s) 1 🟢 -

Say Hi Bot (main-1787230587)

Name Average Improvements Regressions
Scores
Levenshtein 100% (+0pp) - -
Metrics
Llm_calls 0 (+0) - -
Tool_calls 0 (+0) - -
Errors 0 (+0) - -
Llm_errors 0 (+0) - -
Tool_errors 0 (+0) - -
Prompt_tokens 0tok (+0tok) - -
Prompt_cached_tokens 0tok (+0tok) - -
Prompt_cache_creation_tokens 0tok (+0tok) - -
Prompt_cache_creation_5m_tokens 0tok (+0tok) - -
Prompt_cache_creation_1h_tokens 0tok (+0tok) - -
Completion_tokens 0tok (+0tok) - -
Completion_reasoning_tokens 0tok (+0tok) - -
Total_tokens 0tok (+0tok) - -
Duration 1s (0s) 12 🟢 1 🔴

@github-actions

github-actions Bot commented Aug 20, 2026

Copy link
Copy Markdown

Braintrust eval report

Say Hi Bot (main-1787230591)

Name Average Improvements Regressions
Scores
Levenshtein 100% (+0pp) - -
Metrics
Llm_calls 0 (+0) - -
Tool_calls 0 (+0) - -
Errors 0 (+0) - -
Llm_errors 0 (+0) - -
Tool_errors 0 (+0) - -
Prompt_tokens 0tok (+0tok) - -
Prompt_cached_tokens 0tok (+0tok) - -
Prompt_cache_creation_tokens 0tok (+0tok) - -
Prompt_cache_creation_5m_tokens 0tok (+0tok) - -
Prompt_cache_creation_1h_tokens 0tok (+0tok) - -
Completion_tokens 0tok (+0tok) - -
Completion_reasoning_tokens 0tok (+0tok) - -
Total_tokens 0tok (+0tok) - -
Duration 1s (+0s) - 19 🔴

@AbhiPrasad
Abhijeet Prasad (AbhiPrasad) deleted the abhi-feat-sdk-196-filter-pr-metrics branch September 1, 2026 21:28
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants