Add opt-in OS settings inheritance for native themes - #5927
Conversation
|
@codex review |
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
Native fidelity (Android, Material 3)54 pairs compared -- median 95.6%, worst 91.3% ( Distribution --
Geometry vs native (bbox offset / size ratio / center offset / corner radius) -- gated separately from the visual score
Side-by-side comparisons (worst first)
|
|
Compared 172 screenshots: 172 matched. Benchmark ResultsDetailed Performance Metrics
|
Cloudflare Preview
|
|
Compared 157 screenshots: 157 matched. Native Android coverage
✅ Native Android screenshot tests passed. Native Android coverage
Benchmark ResultsDetailed Performance Metrics
|
|
Compared 193 screenshots: 193 matched. |
|
Compared 172 screenshots: 172 matched. Benchmark ResultsDetailed Performance Metrics
ParparVM vs HotSpot (JDK 25): Windows x64Runner CPU: Intel64 Family 6 Model 173 Stepping 1, GenuineIntel (baseline Ratios are ParparVM / JDK 25: below 1.00x ParparVM is faster (time) or smaller (RAM). Median of 5 interleaved, paired rounds; every run's output was verified. A regression is a ratio more than 15% (time) / 15% (RAM) above its baseline in
Result: no regression |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: c16112b318
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
✅ Continuous Quality ReportTest & Coverage
Static Analysis
Generated automatically by the PR CI workflow. |
|
Compared 12 screenshots: 12 matched. |
|
Developer Guide build artifacts are available for download from this workflow run:
Developer Guide quality checks: |
|
@codex review |
|
Compared 172 screenshots: 172 matched. Benchmark ResultsDetailed Performance Metrics
ParparVM vs HotSpot (JDK 25): Windows arm64Runner CPU: ARMv8 (64-bit) Family 8 Model D84 Revision 1, MICROSOFT CORPORATION (baseline Ratios are ParparVM / JDK 25: below 1.00x ParparVM is faster (time) or smaller (RAM). Median of 5 interleaved, paired rounds; every run's output was verified. A regression is a ratio more than 15% (time) / 15% (RAM) above its baseline in
Result: no regression |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 4599d362ef
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
Compared 172 screenshots: 172 matched. ParparVM vs HotSpot (JDK 25): Linux x64Runner CPU: AMD EPYC 7763 64-Core Processor (baseline Ratios are ParparVM / JDK 25: below 1.00x ParparVM is faster (time) or smaller (RAM). Median of 5 interleaved, paired rounds; every run's output was verified. A regression is a ratio more than 15% (time) / 15% (RAM) above its baseline in
Result: no regression |
|
Compared 172 screenshots: 172 matched. ParparVM vs HotSpot (JDK 25): Linux arm64Runner CPU: Neoverse-N2 (baseline Ratios are ParparVM / JDK 25: below 1.00x ParparVM is faster (time) or smaller (RAM). Median of 5 interleaved, paired rounds; every run's output was verified. A regression is a ratio more than 15% (time) / 15% (RAM) above its baseline in
Result: no regression |
Native fidelity (ios-27-metal)68 pairs compared -- median 94.6%, worst 74.5% ( Distribution --
Geometry vs native (bbox offset / size ratio / center offset / corner radius) -- gated separately from the visual score
Side-by-side comparisons (worst first)
|
|
Compared 166 screenshots: 166 matched. Benchmark Results
Detailed Performance Metrics
ParparVM vs HotSpot (JDK 25): macOS arm64Runner CPU: Apple M1 (Virtual) (baseline Ratios are ParparVM / JDK 25: below 1.00x ParparVM is faster (time) or smaller (RAM). Median of 5 interleaved, paired rounds; every run's output was verified. A regression is a ratio more than 15% (time) / 15% (RAM) above its baseline in
Result: no regression |
|
Compared 155 screenshots: 155 matched. Benchmark Results
Build and Run Timing
Detailed Performance Metrics
|
|
Compared 154 screenshots: 154 matched. Benchmark Results
Detailed Performance Metrics
|
|
Compared 150 screenshots: 150 matched. |
|
Compared 155 screenshots: 155 matched. Benchmark Results
Build and Run Timing
Detailed Performance Metrics
|
|
Compared 223 screenshots: 223 matched. |
Add the missing AMD EPYC 9V45 baseline using the completed five-round measurements from run 36737294474, job 109964379342, through the existing calibration script. Preserve every existing CPU baseline and global threshold. Allow the iOS browser typing step to include XCUITest's preceding-action idle wait. Run 36737295337 spent 28 seconds in that wait, exhausting the old 30-second app deadline before the first focus tap. Keep real input-event assertions and the driver's bounded 25-second key loop unchanged. Validation: 23 performance gate tests; all 26 new CPU metrics resolve and pass; Java 8 compilation of every gesture-suite class; git diff --check. The separate CoreSimulatorService hang in run 36737294450 is a runner fault.
|
@codex review |
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 3e8e6e3111
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 54e5ec1db8
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
✅ ByteCodeTranslator Quality ReportTest & Coverage
Benchmark Results
Static Analysis
Generated automatically by the PR CI workflow. |
Native fidelity (iOS Modern, Metal)68 pairs compared -- median 95.0%, worst 83.5% ( Distribution --
Geometry vs native (bbox offset / size ratio / center offset / corner radius) -- gated separately from the visual score
Side-by-side comparisons (worst first)
|
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: ad4b66ecc5
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
@codex review |
|
Codex Review: Didn't find any major issues. Keep it up! Reviewed commit: ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
…runs The same CPU model and unchanged VM/benchmark sources measured allocation ratios of 2.712, 2.595, 3.425, 2.808, and 3.571 across completed runs 36727183244, 36737294581, 36813957271, 36823098403, and 36830698080. Use calibrate-perf-baseline.py on those five perf-results.json artifacts, retaining only this CPU's objectAllocation row. The run medians produce a 2.808 time baseline with 45% row-specific time tolerance under the existing spread-margin policy, and a 0.454 memory baseline. Other rows and global thresholds are unchanged. All five observed runs pass; values above the calibrated band still fail. Validation: 23 performance-gate tests; exact preservation of other baseline rows; identical VM sources across all five revisions.
#5927 added a Windows x64 Intel Family 6 Model 173 calibration and re-measured the Windows ARM64 d49 objectAllocation row in the old single file. base/ is regenerated from master's file, asserted lossless; the old file stays deleted. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>




















































































































































































































































































































































































Native themes can now optionally inherit available OS colors and typography through independent, default-off
useNativeColorsBoolanduseNativeFontsBoolconstants and matchingUIManagersetters. Applications retain control over explicit style and palette overrides, and the existing larger-text accessibility setting applies once after font inheritance.This adds settings snapshots and live/resume refresh hooks for Android, iOS/macOS, Windows, Linux, and JavaSE; deterministic simulator controls; regenerated modern native themes; documentation; and regression coverage. Unsupported settings retain bundled defaults.
Validation
verify: 36 tests passed, SpotBugs reported zero bugs/errors. The final formatting cleanup also passed this focused verify run, with zero Checkstyle violations.git diff --checkpassed.Windows, Linux, Android, and iOS device UI behavior still needs platform testing. Native checks were run locally only.