Skip to content

[Temporary] Unblock BCal Luna evaluation - #958

Draft
martinsrui-msft wants to merge 6 commits into
mainfrom
martinsrui-temporary-bcal-luna-run
Draft

martinsrui-msft wants to merge 6 commits into
mainfrom
martinsrui-temporary-bcal-luna-run

Conversation

@martinsrui-msft

@martinsrui-msft martinsrui-msft commented Oct 8, 2026 •

Copy link
Copy Markdown
Collaborator

DO NOT MERGE

This is a temporary, run-only integration branch to unblock BCal NL2AL evaluations with the currently pinned bc-eval[capi]==0.3.14. It must not be merged, used as a base, or used to modify native stack #945 or any existing PR branch/base.

Included fixes

  • Fix Luna LM Checklist request formatting #943 / 7df241bde459b79211acf953b48a47d1c5291171 — adapt outgoing LM Checklist requests for Responses-compatible assistant output_text content and omit Luna temperature.
  • Fix skipped BCal aggregation #944 / d05f8ca450818b69e37d1e873d86fd35702ae493 — make BCal aggregation/completeness resilient so summaries and artifact handling continue after matrix failures.
  • Fix Luna LM Checklist response parsing #948 / 55a8f41b8c9d0f9fa1078e2f3a2300dedd7caf92 — parse nested Luna Responses output, translate supported parameters, and fail explicitly when assistant text is absent.
  • Report agent timeouts without blocking results #949 / 95503da7810f1e4a83eb72cea19b30adf9bc390f — persist/export timeout rows, keep timed-out matrix jobs failed, score timeouts deterministically with test_passed=0, and allow complete results to reach Kusto.
  • 373008d0ee4d11c686304bacb5465d73586d8f6e — make the two existing optional-value truthiness checks explicit for the Ty version used by CI.
  • 02b00632c927cb2a6e61b03c83a3fbb48f4b1e12 — persist classified CAPI HTTP 5xx failures as BCal infrastructure-error rows and upload them to Kusto without a model-quality score.
  • 1bbe3e76d44bbf44cdd6923f6d69899a0d4f68a2 — adapt the temporary fixes to current main's simplified BCal API and stricter Ty checks.

CAPI infrastructure-error behavior

  • Only explicit CAPI HTTP 5xx failures are classified; generic agent errors are unchanged.
  • The BCal matrix entry still exits nonzero and remains visibly failed.
  • A complete result row is persisted and uploaded, so exact completeness can still reach the full dataset count.
  • LM Checklist is not called for the infrastructure row. Normal rows keep test_passed; timeout rows keep test_passed=0; the infrastructure row uses CoreScore=null and preserves the error in Kusto's Error field.
  • Only BCal enables unscored infrastructure rows. Other evaluation workflows retain the existing global core-score behavior.
  • No retries, generic AgentError scoring, bc-eval changes, model changes, or dependency changes are included.

Temporary scope

Upstream bc-eval work is intentionally not consumed here. The durable implementation will be redone after the M365 LLM API uptake lands in bc-eval; this branch should be discarded after the evaluation run.

Validation

  • Focused runtime/workflow/scorer/completeness suite: 206 passed.
  • Full Ruff check and format check passed.
  • Full BC-Bench Ty and package-wide bcbench-core Ty passed.
  • All pre-commit hooks passed.
  • Exact pinned bc-eval==0.3.14 CLI accepted the infrastructure row with null scores and no judge call; its Kusto conversion produced normal CoreScore=1.0, infrastructure CoreScore=null, and preserved Error.
  • Workflow YAML regression tests passed; actionlint reported no new diagnostics beyond the repository's known custom $/.github/... and environment.deployment syntax.
  • git diff --check passed.
  • Both workflow installs remain exactly bc-eval[capi]==0.3.14.

AB#653421

@github-code-quality

github-code-quality Bot commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Code Coverage Overview

Languages: Python

Python / code-coverage/pytest

The overall line coverage in commit 1bbe3e7 in the martinsrui-temporary... branch is 87%. The line coverage in commit 92af447 in the main branch is 86%.

Show a line coverage summary of the most impacted files.
File main 92af447 martinsrui-temporary... 1bbe3e7 +/-
src/bcbench/commands/result.py 93% 89% -4%
src/bcbench/results/display.py 93% 92% -1%
src/bcbench/results/base.py 94% 94% 0%
src/bcbench/com...ands/dataset.py 25% 26% +1%
src/bcbench/com...nds/evaluate.py 77% 78% +1%
src/bcbench/evaluate/base.py 91% 93% +2%
src/bcbench/age...t/bcal/agent.py 86% 89% +3%
src/bcbench/exceptions.py 71% 75% +4%
src/bcbench/age..._capi_bridge.py 53% 64% +11%
src/bcbench/res...completeness.py 0% 96% +96%

Updated October 08, 2026 12:14 UTC

from bcbench.config import get_config
from bcbench.dataset import BaseDatasetEntry, BugFixEntry, NL2ALEntry
from bcbench.evaluate.base import AgentRunner, EvaluationPipeline
from bcbench.evaluate.base import AgentRunner, EvaluationOutcome, EvaluationPipeline
martinsrui-msft and others added 6 commits October 8, 2026 14:09
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants