Skip to content

Run Flutter widget code on Codename One - #5883

Open
shai-almog wants to merge 409 commits into
masterfrom
flutter-dart-transpilation
Open

shai-almog wants to merge 409 commits into
masterfrom
flutter-dart-transpilation

Conversation

@shai-almog

Copy link
Copy Markdown
Collaborator

Compiles Dart widget source to Java at build time and runs it on Codename One's own renderer. There is no Dart VM in the result, no embedded engine and no platform view — a Dart screen becomes ordinary Codename One components, so it inherits the theme, the event thread, the accessibility tree and the native build.

Two entry points, which are the two things a user actually wants to do:

  • FlutterUI.wrap(widget) returns a Container that goes anywhere an ordinary component does, for putting one Dart screen inside an existing app. It deliberately does not install the Material base theme, so it will not restyle the screen around it.
  • FlutterUI.runApp(widget) mounts the tree as the whole UI, and does install it.

The build wiring is the archetype's own: the transcode-flutter goal is already bound and is a silent no-op until src/main/flutter exists. No Dart SDK is involved — the transpiler is Java and the Dart is input, never executed.

Benchmark

scripts/flutter-bench builds one application two ways — the Flutter toolchain's release build, and the identical Dart source transpiled by Codename One — and publishes size, start-up and idle memory per platform to the PR and to port status.

Neither application is vendored. prepare.sh takes the gallery from the Flutter SDK CI clones and copies the same 159 files into both trees, so the comparison cannot drift: there is no second copy for an edit to land on. An earlier round of this work compared two different galleries and every number it produced was meaningless. The Codename One side is generated from the shipping archetype with the runtime dependency enabled — the two steps the guide tells a user to take — so a change that breaks the documented wiring breaks the benchmark too.

Three measurement decisions, each because the obvious alternative flattered us:

  • Start-up is a bracket. Flutter's FIRSTCONTENT is a UI-thread callback that runs before that frame is rasterised, while our FIRSTFRAME fires once the form is on screen. Comparing them charges one runtime for rasterising its first screen and not the other — which is what the harness this replaces did. Flutter's figure is reported as a range and the ratio uses the end least favourable to us.
  • Executable code is every Mach-O in the bundle, not the main executable. On iOS a Flutter application's own code is not in the executable at all: Runner is a thin launcher and the Dart image sits in Frameworks/App.framework. Sizing the executable compared our whole runtime against their stub.
  • Assets are staged from Flutter's own built bundle. Staging the whole asset package ships files Flutter tree-shakes away; staging only the 1x images ships fewer than Flutter does.

And two refusals rather than a substituted number:

  • iOS start-up and memory are not measured. Dart cannot AOT-compile for the simulator, so a simulator run would time Flutter's JIT debug engine against our release build. Sizes come from release device bundles, which need no signing.
  • Desktop is the native ports only — AppKit, clang-cl, GTK3/Cairo — never JavaSE, whose bundled JVM and dependency directory are not the same kind of artifact as Flutter's native bundle.

Nothing contacts the build server: the *-source and local-* targets build on the runner. Only our own numbers are gated, so a Flutter SDK upgrade that grows their build cannot turn ours red.

Fidelity work in this PR

  • RoundBorder draws itself when there is no shadow and no uiid mode, instead of going through a component-sized offscreen image. The cache hides the cost for a static shape; for one that animates its size it does not — a circle growing to 1618×1618 threw away a 10MB surface per frame.
  • A paint-only shift needs a box it can paint into. Component.paintInternalImpl clips every component to its own rectangle, so FractionalTranslation drew most of itself into the discarded region — the feature-discovery circle rendered as a quadrant with two straight edges meeting at its centre.
  • FittedBox was a pass-through. Flutter lays its child out unbounded and scales the result, which is why text inside one shrinks instead of wrapping.
  • Card rebuilt its border every time, and RoundRectBorder caches its shadow against the border instance, so every rebuild re-rendered it with a gaussian blur. With the cache off it also translates the live Graphics by the shadow offset and never undoes it, so cards drifted 9px down and 4px across, accumulating down the page.

Verification

  • 398 flutter-runtime tests, plus 21 new tests for the benchmark's arithmetic
  • Parity sweep 48/48 routes: device mean 2.06%, median 1.66%; /demo/card 14.98% → 2.30%, /demo/grid-lists 26.30% → 0.32%
  • prepare.sh verified end to end; the generated project builds 563 Java files from 159 Dart files, which also exercises the documented user path
  • actionlint 74/74, copyright and control-character gates clean against the merge base

Not yet proven

No benchmark platform adapter has been exercised end to end — run_bench.py --list reports that rather than implying otherwise, and the first CI run is the thing to review. The native compile steps (xcodebuild, gradle, clang-cl, GTK3) have not run on a runner.

This is also the first mention of the feature in docs/, which has carried none until now.

shai-almog and others added 30 commits September 18, 2026 21:41
…ntains

A Scaffold is a Material, and a Material is what sets the default text style
for its subtree. Ours set none, and nothing else did either under a nested
theme: a Theme is an inherited widget and wraps nothing, so installing one
changes what Theme.of answers while the ambient text style stays exactly as
whoever built it last left it -- the application's.

The effect is not that a page ignores its theme, which would have been noticed.
It is that a page takes from its own theme only the fields its styles actually
set, and silently keeps the application's for the rest, because a Text merges
its own style OVER the ambient one. So sizes came from the right theme and
family and line height came from the wrong one.

The typography demo is the clean case, since it renders the type scale itself.
Its 96sp display role states no height and no family, so it inherited the
application's Montserrat at a line height of 1.43 -- a body role's height
applied to a display role. Its two wrapped lines sat 141 logical pixels apart
against the reference's 115, and every item on the page compounded the error,
so the further down the page, the further out of register it went.

/demo/typography: 9.76% -> 7.52% wrong pixels, which takes it off the
over-target list; the sweep mean goes 2.67% -> 2.65% and the routes above 8%
drop from two to one.

Recorded honestly: /demo/motion goes the other way, 4.35% -> 6.21%. Its text
now resolves against the demo's theme rather than the application's, which is
the correct answer and measures worse, so something below it is reading the
wrong role -- worth a look on its own rather than a reason to keep the ambient
style wrong everywhere else.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A list tile's height comes from how many LINES it has. The leading and trailing
widgets are centred inside that height and are allowed to overflow it; they
never drive it. Ours took the tallest of text, leading and trailing, so a tile
was as tall as whatever was in it.

Crane's destination list is where it showed. Its rows carry a 60lp thumbnail,
which with the vertical padding came to 76lp against Material's two-line height
of 72lp. Four logical pixels a row does not look like anything on one row, but
the rows below it each start 4lp lower than the row before, so the list drifts
steadily out of register with the reference -- 43 device pixels by the fourth
row, which is most of a line of text.

Two-line height (72lp) was also simply missing: every tile used the one-line
minimum of 56lp, so any tile whose content happened to fit was the wrong height
even without a leading widget. The existing test asserted 56lp for a tile it
built WITH a subtitle, which is why nothing caught it; updated to 72lp, with
the child offsets that follow from it.

The text may still push a tile past its nominal height -- that part Material
does allow -- so this keeps the max against the text block and drops it only
for the leading and trailing.

/crane 8.69% -> 7.64% wrong pixels, /demo/motion 6.21% -> below 4.6%, sweep
mean 2.65% -> 2.58%, and no route in the gallery is over 8% any more.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…g child grow

Two things were missing from the scaled variant, and together they were the
worst single frame in the motion suite.

Every leg of this pattern is EASED, on three different curves: the arriving
content decelerates in (0, 0, 0.2, 1), the leaving content accelerates out
(0.4, 0, 1, 1), and both scales run on the standard curve (0.4, 0, 0.2, 1).
All three ran linearly here. Linear is not a small difference on a fade that
only occupies the first three tenths of the run: a third of the way through
that window the curve has taken the outgoing surface most of the way out while
linear has barely started, so the frame still looks like nothing has happened.

The outgoing child also never scaled. It should keep growing, through 100% to
110%, while the incoming one comes up from 80% -- that continuing motion is
what makes the two read as one surface handing over rather than as a
cross-fade. Ours pinned it at 100% for the whole run, because the scale was
computed from the incoming animation only, which for the outgoing child sits
at 1 the entire time.

reply_search: worst frame 21.77% -> 4.42% wrong pixels, mean 5.96% -> 4.00%.
It was the worst step in the suite and is now the best.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Same defect as the shared axis, in the pattern beside it: the fade-through
leaves on one curve (0.4, 0, 1, 1) and arrives on another (0, 0, 0.2, 1), and
both legs ran linearly.

The old page is the one you watch here, since it carries the whole screen for
the first three tenths while the new one is still invisible: a third of the way
into that window the curve has it 58% gone where linear has it at 44%.

Measured no change on the suite, and the reason is worth recording rather than
leaving for the next person to re-derive: the step named "search" exercises the
SHARED AXIS transition, not this one. This is the Reply body's own switcher,
which no step drives yet. Corrected on the strength of the pattern's own
specification.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The platform page push ran the incoming page on linear-to-ease-out. That is the
curve of the page being LEFT, and of the shadow; the page arriving rides
fast-ease-in-to-slow-ease-out, which is a three-point cubic -- two beziers
joined at a point, accelerating hard to half its travel in the first fifth of
the run and then changing character and settling slowly.

A single cubic cannot express that, so Motion gained
createThreePointCubicMotion. Each segment is solved in its own normalized space
and rescaled, which is what makes the halves meet exactly at the joint rather
than step there.

This was hard to see because both curves start at 0, and both finish at 1 at
exactly 500ms: the error is invisible at either end and worst in the middle.
Measured against the reference as a fraction of travel completed, where our own
motion tracked linear-to-ease-out to within 0.3ms at every sample:

    ms     reference   three-point   linearToEaseOut
    50        0.2382        0.2383            0.2615
   150        0.7600        0.7604            0.7195
   250        0.9422        0.9422            0.9201

rms error over the run: 0.0005 against 0.0266. In pixels that is a whole page
sitting 46 device pixels from where it belongs halfway through every push.

push_reply mean 8.84% -> 5.81% wrong pixels (worst 13.08% -> 10.69%),
push_shrine 5.68% -> 5.14%, push_demo_app_bar worst 11.62% -> 11.16%. Every
step in the motion suite is now under the 8% target.

Known gap, recorded rather than papered over: the outgoing page should ride
linear-to-ease-out while the incoming one rides this, but a Codename One slide
drives both pages from a single Motion, so both share this curve. The outgoing
page travels a third of the distance, so the residual there is a third the size.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The platform push was a slide, and a slide is the one thing this transition is
not. A slide holds the two pages a fixed screen apart and moves the pair, so
they travel as one rigid object. On this platform the arriving page crosses the
WHOLE screen while the page it covers drifts only a THIRD of it, each on its
own curve, and the gap between them closing as they travel is exactly what says
one is in front of the other.

Moving both the full distance is not a subtle error. A third of the way through
a push, the strip of the old page still showing is not dimmer or slightly
shifted -- it is a different PART of that page. Measured at 150ms of a 500ms
push on a 1125px screen, by correlating each side's strip against its own
pre-transition frame:

    ours   old page 855px to the left   (full width * three-point curve)
    ref    old page 270px to the left   (width/3  * linearToEaseOut)

Both figures land on their formula exactly, which is what identifies the defect
rather than merely measuring it.

CommonTransitions cannot express this -- one Motion drives one offset -- so
this is a transition of its own, alongside the container transform. It buffers
both pages and places them independently, and copy(reverse) swaps which one
crosses the screen and which drifts home.

    push_reply        mean 5.81% -> 3.47% wrong pixels, worst 10.69% -> 4.12%
    push_shrine       mean 5.14% -> 2.83%, worst 10.58% -> 3.33%
    push_demo_app_bar mean 7.54% -> 5.24%, worst 11.16% -> 6.42%

Against where the pushes started, push_reply was 9.50% mean and 13.37% worst.

Known gap: the reverse plays on the forward curves rather than the flipped
ones the platform specifies. No step pops a route yet, so that is unmeasured
here and left rather than guessed at.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
InputDecoration.collapsed kept only the hint text and dropped everything that
makes it collapsed, so the result was indistinguishable from a plain
InputDecoration and every such field still drew the full outlined box. No
border and zero padding are the POINT of the factory, not defaults it happens
to inherit.

Reply's compose screen is where it showed: its subject line and its message
body both sat in outlined rectangles where the reference has neither -- the
reference draws only the text, with the section dividers below supplying the
horizontal rules.

The styling arguments the factory already accepted were also being discarded
rather than applied, so an explicit border, hint style, fill or fill colour
passed to collapsed() did nothing. They are applied now, with the explicit
border still able to override the none default.

reply_compose settled frames 5.82% -> 5.32% wrong pixels, mean 7.17% -> 6.76%.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…lose

Two defects a whole-screen percentage could not see, both found by using the
app rather than by reading a number.

THE SHADOW. paintShadowRings stacks four FILLED shapes, so each ring darkens
everything inside it as well as its own band, and all four used one alpha. The
band against the surface was therefore hit four times and came out about twice
as dark as Material's shadow, with a linear falloff where Material's is a blur.
On the Reply study's compose button -- a 56dp circle, the largest elevated
surface in the gallery -- that reads as a grey box around the disc, because the
rings follow the surface's bounding geometry and the accumulated alpha is high
enough to see. Calibrated per ring against the reference:

    band   reference   one alpha for all   these alphas
       1       0.071               0.148          0.071
       2       0.035               0.077          0.035
       3       0.020               0.039          0.020
       4       0.008               0.039          0.008

THE CLOSE. ContainerTransformTransition used its `closing` flag for one thing:
picking a mirrored curve. Everything else -- the geometry, the scrim, the
colour, the crossfade -- played the OPENING animation regardless of direction,
and the origin was looked up on the source form, which on the way back is the
page being left rather than the page being returned to. So it found nothing,
fell through to its no-origin guess, and grew the page out of the middle of the
screen a second time instead of folding it back into the button.

Now the direction is carried through properly: the anchor is found on whichever
page is NOT travelling, the buffers swap roles, the geometry runs backwards,
the colours cross the other way, and the scrim lifts smoothly across the run
in proportion to what is left rather than mirroring the opening's fifths --
mirroring held it at full black through the middle and dropped it in one step.

The mirrored curve is gone with it: 1 - curve(elapsed) already IS the mirrored
easing, so selecting a mirrored curve as well gave curve(1 - elapsed), which is
a different motion.

A pop is now MEASURED. Every step in the suite was an arrival, so a back
transition could be completely wrong with every number still green -- which is
exactly what happened. The new compose_back step scored 96.90% worst and 29.68%
mean on the code as it stood; it is 13.13% and 5.26% now.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A button consumes its content rather than mounting it, so it walks down to the
Text or Icon it can actually render. That walk stops at any wrapper it cannot
open, and it could open exactly one of them -- Tooltip. Padding, Center, Align,
SizedBox, Container and Expanded were all opaque to it, so a button whose label
sat inside any of them rendered NO label at all.

This is not a corner case. Shrine's login buttons are written the ordinary way,
`child: Padding(child: Text(...))`, and both rendered as blank shapes: the row
that should read CANCEL and NEXT was a small pink blob. Reply's compose screen
lost a whole row the same way -- its sender address is a PopupMenuButton whose
child is a Padding around a Row, so the account line was simply absent.

The runtime was saying so on every build: "child Padding is neither a Text nor
an Icon and could not be resolved to one; the button renders no label",
repeated for ElevatedButton, TextButton, IconButton and PopupMenuButton. Those
warnings are now gone -- the count in a full sweep of the gallery is zero.

Implemented by having those widgets declare HasChild, which is what the walk
already uses; each of them already exposed getChild(). Nothing else changes:
the elements that mount these widgets keep laying their child out themselves,
and the two render elements that cast their own widget to HasChild are
unaffected.

/shrine goes 2.68% -> 2.87% wrong pixels, and that is the metric being wrong
rather than the screen: rendering two buttons that were previously absent adds
pixels, and they are still mispositioned (the OverflowBar ignores its end
alignment) and CANCEL takes the wrong colour. Both are now visible problems
with visible buttons, which is the better place to be.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A Codename One theme gives a text field one border unselected and a different
one -- here, none -- selected. That is a reasonable default for a native-looking
field and wrong for this one: a reference text field keeps its decoration
whether it has focus or not, and only the caret and the highlight colour
change. We set a border explicitly only when the decoration names one, so
everywhere the border came from the theme, the focused field simply lost it.

What that looks like in use is worse than a missing outline: the box appears to
belong to whichever field is NOT being used, and it jumps to another field as
you touch things. Measured on Shrine's login screen, the outline's top and
bottom edges:

    on arrival           y = 1191, 1363      (the password field only)
    after tapping into
    the password field   y =  979, 1151      (the username field only)

Both fields now keep their outline, and tapping between them changes nothing:
four edges, at 979, 1151, 1191 and 1363, before and after.

/shrine goes 2.87% -> 3.17% wrong pixels. As with the buttons, that is the
metric rather than the screen: a box that was missing is now drawn, and its
corner radius and colour do not yet match -- Shrine's fields are beveled and
ours are rounded. A wrong-looking box in the right place is a smaller problem
than no box at all, and a visible one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A screen you can enter and cannot leave, found by using the app. Both of the
ways out of the gallery's splash page were dead.

IGNOREPOINTER IGNORED NOTHING. It was a structural pass-through: it rendered
its child and left that child fully live. The widget exists for exactly one
reason -- so that something ABOVE it receives a touch landing on top of its
child -- and with the child still answering, the handler above never runs. The
splash slides the home page down, leaves a strip of it showing, and wraps that
strip in an IgnorePointer inside a detector whose tap dismisses the splash. The
strip's own rows took the tap instead, so the dismiss never fired.

Codename One's hit test walks UP from the component it lands on for as long as
each one ignores pointer events, so the flag has to reach the whole subtree
rather than its root: a touch landing on a nested child would stop there.

ONVERTICALDRAGEND WAS NEVER INVOKED. The callback was stored and called from
nowhere in the runtime, so every flick gesture in the gallery did nothing. That
is the splash's second exit -- an upward flick on the same strip, which calls
reverse() directly -- and it is also how the splash is meant to be ENTERED.
Velocity is measured over the whole press rather than the last few moves, which
understates a flick that started slowly: that misses a gesture rather than
inventing one, and every caller compares against a threshold.

Deliberately NOT fixed here: Notification.dispatch() is still a no-op, so the
splash cannot be opened by its own gesture even now. Wiring it needs the
listener's type, and NotificationListener<T> erases T -- delivering
notifications without a type token would hand every ScrollNotification to a
listener waiting for something else, so merely scrolling the gallery home would
open the splash. That wants a type token from the transpiler, not a guess in
the runtime.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Eight em dashes had crept into the javadoc of a file this branch adds to the
Maven plugin. Java sources here are ASCII only -- the toolchains that compile
this tree read them as ASCII, and a single non-ASCII byte fails the build with
"unmappable character for encoding ASCII" in a file nobody was editing.

Replaced with the ASCII double dash the rest of the tree uses. No behaviour
change; the plugin still builds.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A measured inventory of the remaining gaps, kept in the module rather than
under docs/, which is published.

The pixel metrics cannot see the largest one. A screen can render at 2% wrong
and do nothing when it is touched, so the sweep reports the gallery at 2.56%
mean while twenty-one callbacks the gallery actually passes are stored by their
widget and read by nothing -- swipe-to-dismiss, route generation, form
submission, every Cupertino picker. Four such defects were found by using the
app for two minutes while every number stayed green.

Each entry is measured rather than estimated, and says what closing it
requires. Two censuses back it: callbacks that nothing reads, and parameters
that nothing reads, each cross-referenced against what the gallery's generated
code actually sets so a latent gap is not confused with a live one. The
parameter census is per FIELD rather than per file -- Visibility.visible and
ClipRRect.borderRadius are honoured and are deliberately absent, where an
earlier file-level pass wrongly flagged them.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The census counts call sites and cannot tell a broken feature from one something
else compensates for. Two checked by driving the running simulator:

onGenerateRoute reads as the biggest gap in the table -- ten uses, and it is the
gallery's entire route factory -- and is not a live defect at all: Navigator
resolves named routes from its own table, and tapping through a category into a
demo navigates and renders with zero errors. Recorded as a latent risk instead.

Dismissible is the opposite. A full left swipe across a mail card changes 0.0%
of the list: swipe-to-dismiss does nothing whatsoever, where the reference
dismisses the mail. That is the first thing to fix.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Dismissible rendered its child and dropped everything else, so a full swipe
across a mail row in the Reply study changed 0.0% of the screen. It now tracks
the drag, shows the background behind the moving row, honours the per-direction
thresholds, asks confirmDismiss, and fires onDismissed.

Three things had to be got right, each of which failed silently first:

WHO HEARS THE GESTURE. The rows live in a vertically scrolling list, and
Codename One decides who owns a gesture before a component sees it -- a
horizontal drag inside a vertical scroller goes to the scroller, so the row is
never asked. This listens on the FORM, which runs before that decision, and
consumes the event; that is the same hook Tabs uses to swipe between tabs. The
axis is decided once per gesture, past the touch slop, and never revisited, so
a diagonal drag does not fight itself.

WHERE THE ROW IS. Neither end of the tree answers that. A render element inside
a scrolling list reports (0,0) for its position -- placement lives on the
components -- and the first component going down is a leaf, a label inside the
card rather than the card. Both were tried; both hit-tested the wrong rows. The
bounds are the union of every component beneath the row.

HOW THE ROW MOVES. Re-placing the child from the element's own origin puts it
in the corner, for the same reason. It is moved by the DIFFERENCE since the
last frame, applied to the outermost components only -- a component's x is
relative to its parent, so shifting a nested one moves it twice.

Verified by driving the simulator: a left swipe stars the mail and springs back
(Reply's confirmDismiss vetoes the dismissal outside the starred mailbox, and
that veto is honoured), a right swipe past 0.8 of the width deletes it, and the
correct mail is the one that goes.

KNOWN, AND THE NEXT THING TO FIX: deleting an item exposes a defect in the
virtualised ListView. The rows that move up render their new content over their
old, and a provider lookup in a rebuilt row answers null. It survives a
revalidate, a full repaint and scrolling away and back, so it is two live
component sets rather than stale pixels. It is not this widget's doing -- the
list was simply never asked to lose an item before, because nothing could
remove one. Recorded in READINESS.md.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… blocks

Implementing swipe-to-dismiss made removing a list item possible for the first
time, and the first removal showed that a virtualised list does not survive one:
every row below the deleted one draws its new content over its old, and a
provider lookup in a rebuilt row answers null.

Recorded with the seven suspects already ruled out by measurement rather than by
reading -- the dismissed row staying mounted, the callback's timing, repaint
ordering, missing keys, ObjectKey equality, the keyed reconciler, and
RenderHost.detach -- so the next person does not repeat them. The remaining
candidate is named as a hypothesis and not as a finding.

It blocks the rest of phase one: every other callback that removes something
would land on the same defect.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
RenderHost.detach removed an unmounted element's component only when its parent
happened to be that host's own container. An element's component does not
always sit there: a scroll view puts its content pane inside its ScrollPane,
and other effects re-parent too. Where the parent was something else the guard
failed silently and an unmounted element kept a live component on screen.

Not the cause of the list-removal defect -- that is recorded in READINESS.md,
and this does not fix it -- but a component whose element has been unmounted
must not remain in the scene under any container, so the guard was wrong on its
own terms.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Traced by identity rather than class name, which was the thing that had been
hiding it: two different panes print the same ancestor path, so the chain
looked shared when it was not.

The finding is that a scroll view ends up with TWO content panes under one
ScrollPane, both populated, with no detach for the first -- so the old content
subtree is never unmounted and both render. The element tree is right and the
Codename One component tree keeps the previous subtree, which is the conflict
between the two trees rather than a fault in either.

Records the nine suspects ruled out by measurement so they are not revisited,
and names the one remaining question: what replaces the scroll view's content
field without deactivating what was there.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…assignment

Three more measurements, and they clear the scroll view rather than convict it.

The duplication is created by the deletion, not pre-existing: a census before
any swipe reports 19 labels and no duplicates. No element is mounted twice.
And the list's scroll element never replaces its content -- the only content
change during a deletion belongs to an inner image-row scroller inside a mail
row, seen for the first time, built from null.

So the second pane is not a replacement subtree. It is attached to the list's
host by another element, which moves the question to which host an element
joins when it mounts: a RenderHost is flat, an effect element opens a nested
host around its own pane, and an element that should have joined the inner one
joining the outer instead puts a second pane beside the first.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Element.visitChildren is EMPTY by default, so an element that owns children and
does not override it owns them invisibly. DismissibleRenderElement owns three --
the row, the background and the secondary background -- and overrode nothing.

The walk that unmounts a subtree is the one that suffers. Deleting a mail
deactivated five Dismissible elements, and not one of their subtrees was
reached: their elements stayed mounted and their components stayed in the
scene. Every surviving row then drew its new content over its old, an orphaned
subtree was rebuilt with its ancestors gone and its provider lookup answered
null, and the inner lists inside those rows were never unmounted at all -- a
counter on them only ever climbed.

How it was found, since the symptom pointed everywhere but here. The element
tree looked right at every level: keys emitted, ObjectKey equality sound, the
keyed reconciler faithful to Flutter's, attachOrder matching the container's
child count exactly, no element mounted twice, no component attached while
already parented, and RenderHost.detach never once skipping its guard. That
last fact was the tell rather than an acquittal: detach was not skipping, it was
never being CALLED. Logging every deactivation showed five Dismissibles going
while the lists inside them never unmounted, which puts the fault between
deactivation and unmount -- and that is visitChildren.

Deleting a mail now leaves the list correct, with no doubling and no errors, and
a left swipe still stars without dismissing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…efect

Two tests on the Dismissible's children, and the readiness entry rewritten from
blocking to resolved with the cause and the route to it.

The bug class is closed rather than the instance: a census of every element that
owns children finds no other that fails to override visitChildren, so this was
the only one, and it was introduced with swipe-to-dismiss.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Verified before implementing, which is what the plan asks for, and the
verification moved the work: the Cupertino picker's callbacks are not the
problem, the tap that would reach them is.

Tapping any row of /demo/cupertino-picker changes 0.0% of the screen and raises
no error. Everything behind the tap is present -- the transpiler emits the
showCupertinoModalPopup call, it delegates to Dialogs.showDialog, and that
presents a real dialog. What fails is the gesture, and only where detectors
nest: the press reaches both the page-covering overlay and the row's, the row's
reports it owns the tap, and then the release reaches neither. The identical
code path on the home screen, which has no outer detector, delivers the release
and fires the tap.

Recorded as blocking the picker family, since onSelectedItemChanged cannot be
reached while the tap that opens the picker never fires.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…s the harness

I recorded a blocking runtime defect that does not exist. The Cupertino picker
demo is not inert: every row opens its picker, and always did.

bench_pointer's steps defaulted to 8, so a tap injected without naming steps
carried eight drag events at the same coordinates. A zero-distance drag has no
dominant axis, so a scrollable ancestor grabs the gesture and
Form.pointerReleased takes its dragged != null path and never delivers to what
was pressed. Instrumenting that method rather than reading it showed
dragged=ScrollPane on the failing screen and dragged=null on the working one,
and a tap with steps: 0 opens the popup and changes 99.5% of the screen.

Every "changed 0.0%" verdict taken with a defaulted tap is void and is marked
so. Two re-checked: the picker works, and the action chip's 0.0% is correct
because the gallery passes it an empty callback. The Dismissible verdict stands,
since that was a genuine multi-step swipe.

What is left is real and smaller: the popup opens but the picker inside it draws
as a narrow pill instead of a wheel.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
showCupertinoModalPopup delegated to showDialog, which centres a dialog packed
to its content. A modal popup is a different shape: it rises from the bottom
edge and spans the screen.

That is not a near miss. A sheet states a height and leaves its width to the
presentation, so packing sized it to its content and the picker demo's 216-high
sheet came up as a tall narrow strip down the middle of the screen. Anchoring it
south fixed the position and left the strip; the width needed showStretched,
since packing sizes BOTH axes.

Dismissible by touching outside it now, as the reference is by default.

The sheet is the right size and in the right place; what it contains is still
empty, because CupertinoDatePicker.build returns a bare Container and
CupertinoPicker returns a plain Column rather than a wheel. Both are recorded in
READINESS.md as the stubs they are -- this change is about the presentation that
was hiding them.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…visible to the sweep

The popup's presentation is fixed; its contents are empty because
CupertinoDatePicker returns a bare Container and CupertinoPicker a plain Column.
Recorded as unimplemented widgets rather than wiring faults, with the note that
the sweep cannot see any of it: it photographs settled routes, the picker rows
render correctly, and the wheel only exists inside the popup.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…e states

Two ways a field ignored the theme around it, both found by putting our screen
beside the reference's rather than by reading a number.

THE TYPE. Only the field's OWN style was applied, so a field that states none --
which is most of them -- fell back to whatever Codename One's default font
happens to be. It resolves the theme's style now, and bodyLarge rather than
titleMedium: the latter is the Material 2 answer, and Material 3 resolves a text
field through _m3InputStyle, which is textTheme.bodyLarge.

THE LABEL COLOUR. inputDecorationTheme.labelStyle was never read, so a theme
that states the colour of its labels did not get it and every field fell back to
the generic hint colour.

The Rally study is where both show. It sets bodyLarge to a 40-point serif and a
label colour to go with it; its login fields came up in small dark sans where
the reference reads them in large light serif. They now match in size, family
and colour.

Recorded honestly: /rally goes 2.98% -> 3.57% wrong pixels. That is the third
time on this screen that a correct fix has raised the number, and the reason is
the same each time -- the residue is the DECORATION, which is still a full
outline where the reference is a filled box with a single underline, and text
that is now large and light disagrees over more pixels than text that was small
and dark. The remaining work is the decoration, not the type.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…a button its label's colour

Two things dropped on the way to the screen, both visible on Shrine's login.

OverflowBar built a bare Row and kept neither of the two things it was given.
Shrine states an END alignment and 8 logical pixels between its buttons, so
CANCEL and NEXT sat hard against the left edge with no gap where the reference
has them apart and to the right.

A button takes its child's STRING and drops the style that came with it, so a
label that states a colour lost it and wore whatever its role gives. Shrine's
CANCEL asks for onSurface and came out in the pale pink of a text button
against the reference's near-black. The label's own colour is the most specific
thing anyone said about it, so it is applied last.

/shrine 3.17% -> 3.14% wrong pixels, and the row now reads the way it should.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…way, and give the fields their real decoration

Six things, found by folding the compose page back into its button and
looking at what the frame actually contained.

Navigator.pop unmounted the popped route's element tree BEFORE showBack()
started the transition, and a transition snapshots its outgoing Form the moment
it starts -- so the container transform shrank a BLANK card back into the
button, where the reference shrinks the page you were reading. The teardown now
waits for the returned-to page's show to complete.

The transform then drew both contents at 1:1 and clipped them, which makes it a
window rather than a transform: the page inside a half-sized box was the page's
top-left quarter. Both are now scaled to the box's width, and the tapped thing's
content fades rather than standing at full strength.

A button that could not reduce its child to a string or a glyph drew NOTHING.
The compose page's account row is a PopupMenuButton whose child is a Row of an
address and a caret, so the row was a blank band -- flutterfan@gmail.com simply
was not on the page. A button now mounts a child it cannot consume and makes it
transparent to touch so the press still lands on the button underneath.

Divider read only its own parameters: not the ambient DividerTheme, not
colorScheme.outlineVariant, and not indent/endIndent. outlineVariant falls back
to onBackground as Flutter's does, which is how a hand-written scheme gets the
near-black rules the compose page draws.

Chip was a bare Row -- an avatar and some text loose on the page, with no pill
behind them. It has its capsule, its background and its label padding.

A field with no border of its own kept the Codename One theme's box, where
Flutter's default is a rule underneath. And a translucent ink was painted at
full strength, because a CN1 Style foreground carries no alpha: the compose
page's Subject placeholder asks for the primary colour at half opacity and read
as a title.

48 routes: mean 2.57% -> 2.50% wrong pixels, shrine 3.14 -> 2.47, text-field
6.13 -> 5.12, rally 3.57 -> 2.94, cupertino-text-field 2.95 -> 2.29, nothing
regressed, 0 errors. 394 tests green.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…what the surface carries

Four faults, each found by fitting the reference's own frames rather than by
reading the spec.

The painted surface is NOT the rectangle. Material lerps the tapped thing's
outline into the page's, and a circle squashes the rectangle toward a square
about its centre as it goes -- that is what keeps a round button round instead
of stretching into a lozenge the instant it starts to grow. Fitting the
reference's middle frame: the rectangle is 685 x 1392 and the painted surface
685 x 1066, same box, same centre, 326 pixels shorter. Painting the rectangle
put a third of the card's height in the wrong place for the whole run, and it
explained every pixel of the vertical disagreement -- the horizontal already
matched to within one pixel.

The rounded-rect path swept its corners the wrong way. GeneralPath.arcTo is
counter-clockwise by default and the rectangle is walked clockwise, so every
corner took the 270 degree route and bulged back into the box. Harmless-looking
at a small radius; this transition's corners reach half the short side.

The tapped thing's content was painted from its COMPONENT, and what is drawn on
top of a surface can be a separate component beside it rather than a child of
it -- so the compose button's pencil was absent from the card for the entire
transform. It is now cut out of the page's own photograph, masked to the
button's outline and to what is not the button's own colour, which is exactly
its content. The surface colour comes from the commonest colour in that
cut-out rather than the middle pixel, which was usually the glyph.

And the closed content does not fade. The fade variant holds it at full opacity
and simply covers it as the page arrives; I had it fading, which is a dissolve
rather than a transformation. Measured: the pencil inside the reference's
shrinking card is pure black, not a tint.

Motion suite 8/8 within ratchet. The two container transforms:
reply_compose worst 12.48% -> 5.82%, mean 7.04% -> 5.07%;
compose_back  worst 13.36% -> 5.68%, mean 5.71% -> 4.22%.
Static sweep unchanged at 2.50% mean, 0 routes above 8%, 0 errors.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…tBugs gate

Image.scaled hands the work to the platform, and two of the three platforms that
matter answer with a POINT SAMPLER: the desktop port reads one source pixel per
destination pixel and discards the rest, and Android asks createScaledBitmap not
to filter. Enlarging, that is merely blocky. Reducing, it is aliasing -- jagged
edges and moire on anything with detail in it, which is every downscaled
photograph and every icon drawn smaller than it was authored.

Image.scaledSmooth averages the whole source rectangle behind each destination
pixel instead, and fillSmooth is the same for a centre crop. Both are ADDITIVE:
scaled() and fill() are untouched, so no existing rendering moves and no
screenshot baseline has to be reseeded. The runtime's image widget uses them
wherever the picture gets smaller, where the result is cached anyway.

The fractional edge weights are the whole of it. Rounding each box to whole
source pixels makes consecutive boxes one and two source pixels wide in turn,
which moves their centres off the grid they are sampling: measured, that alone
shifted a downscaled photograph a full pixel and cost more than the filtering
gained (11.62% wrong pixels against the point sampler's 7.74%). With the weights
the same route measures 5.90%, and it beats a bilinear Java2D step in the port
as well, which scored 6.21%.

The grid-list demo is now 5.90% from 7.74% and the sweep's mean 2.50% -> 2.46%,
with no route moving the wrong way. Every remaining pixel there is photo detail:
the tile geometry is identical to within one pixel on both axes and the two
screens are indistinguishable.

Also clears four SpotBugs findings, which is a ZERO-findings gate: two are from
the container transform above (a boxed Integer constructor, a null check the
analyser can prove redundant) and two pre-date it -- CupertinoPageTransition
dereferences a motion it guards only through its sibling. Both motions are set
and cleared together, and now the guard says so. core-unittests: 0 findings.

Motion suite 8/8 within ratchet, static sweep 48/48, 0 errors.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
shai-almog and others added 30 commits September 29, 2026 14:39
The committed chart-time reference labelled every time-axis tick "Invalid
Date": a Java long reached the browser's Date as a boxed object and did not
convert. With longs that fit in 53 bits held as plain numbers (da0e534),
the axis now reads Mar 2 .. Mar 8, 2024 -- the golden recorded the bug.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The hosted Linux x64 pool now also hands out AMD EPYC 9V45 runners, which
had no rows, so the gate failed on any job that landed on one. Calibrated
from that job's perf-results.json as calibrate-perf-baseline.py documents;
only the new model's rows are added. On the runners that do have rows, the
same head measured within tolerance of the baseline on every benchmark.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The function-name minifier protects any generated function whose name
extends a quoted cn1_ string of 16+ characters, so a name the bridge JS
builds by concatenation (the screenshot runner's lambda stem in port.js)
still resolves. It took those stems from the bundle's own string literals
too -- and the bundle quotes every instance field's property name in its
class f: lists. A field strokeWidth therefore protected strokeWidth_R_double
and strokeWidth_double, the getter and setter every transpiled Dart property
has: 5,211 functions of the transpiled Flutter gallery kept their full
names.

Stems now come from the bridge sources only. The bundle's literals are
complete names (jvm.setMain's main method, the field properties) and keep
their exact-match protection.

Transpiled Flutter gallery, same translator input, script bytes: -151KB
(-1.2%) with this alone. vm/tests JavaScript group: 374 run, 0 failures.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
f377038 gave every class translated code allocates with new its own
allocation function. On the transpiled Flutter gallery that was 3,153
functions and +2.2MB (+17%) of script, which the javascript benchmark leg
gates at zero tolerance against its baseline.

An allocation function pays off where V8 sees the allocation many times:
- a new inside a loop -- between a backward branch and its target (jumps
  and switches) -- which is where V8 removes one that never escapes;
- the boxed types, always: their instances are built in their own valueOf
  factories, never in the loop that autoboxes, and hashMapChurn spent 21%
  of its time building Integers through jvm.newObject without them.
Every other new -- a listener, a UI component, a lambda built once per call
-- uses jvm.newObject again. A class with more than 16 instance fields never
gets one: the literal lists every inherited field, and 63 wide classes were
a third of the functions' size. The per-class classDef constant is numbered
rather than named after the class, which the minifier never shortened.

Tried and dropped: treating new in any method of <= 16 instructions as hot
(to catch small factories): +490KB and it did not reach valueOfHeap.

Transpiled Flutter gallery, same translator input, against the translator
before this series: script 12,855,430 -> 12,705,943 bytes (-1.16%, with the
previous commit), deflated download -0.03%. Compute workloads under node,
production-mode translation: objectAllocation 24ms, valueEscape 46ms,
hashMapChurn 236ms, stringBuilding 234ms -- unchanged from the previous
head; checksums identical to a JVM. vm/tests JavaScript group: 374 run,
0 failures.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ng it

The runtime bound System.<clinit> to a native that only created the console
print streams, so every other static of java.lang.System was never
assigned -- LOCK above all, which System.gc() and the GC thread synchronize
on. Any app that reached System.gc() (the Codename One core does, from the
EDT) died on "Cannot read properties of undefined (reading '__monitor')";
the transpiled Flutter gallery did, right after its first frame, showing an
internal application error dialog.

The binding now runs the translated initializer (installNativeBindings keeps
it in jvm.translatedMethods) and then installs the console print streams as
before. It stays a plain function so System's initializer, and every
class-init guard in front of System, does not become suspending; a
translated initializer that returns a generator is handed to the class-init
driver with the stream swap after it.

Verified on the transpiled Flutter gallery in headless Chromium: the
__monitor error and the dialog are gone, with the old and the new translator
alike (the defect predates this series). vm/tests JavaScript group: 374 run,
0 failures.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
ParparVMBootstrap called the lifecycle's init() and start() inline on the
worker's main thread, where every other port runs them on the EDT (iOS
hands its stub to Display.init). Code that checks threads broke: a
transpiled Flutter app's runApp asserts it is on the EDT, so the gallery
threw IllegalStateException("Flutter framework mutation off the EDT") and
never left its loading screen, with the old translator as with the new.

The bootstrap now queues itself with callSerially after Display.init and
the afterInit hook, so the hardening stamp still runs before init/start.

Verified: the transpiled Flutter gallery reaches its first frame in
headless Chromium (BENCH:FIRSTFRAME ~760ms) where it used to stop at the
loading screen; vm/tests JavaScript group 374 run, 0 failures. NOT run
locally: the JavaScript screenshot suite (scripts-javascript.yml), whose
build uses this bootstrap and is the gate for this change -- it needs the
hellocodenameone build through ~/.m2, which this checkout does not use.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…rom new

In the transpiled Flutter gallery -- the build the javascript benchmark leg
measures -- hashMapChurn ran 1488ms in headless Chrome against 227ms for the
same workload in a small application. The suspension fixpoint had made the
whole java.util.HashMap API a generator: Vector, Hashtable, SimpleTimeZone and
the Collections.synchronized* wrappers have synchronized hashCode/equals, so
every hashCode()/equals() dispatch site that could reach them suspended, and
through computeHashCode/areEqualKeys every map operation and every caller of
a map did too. A small application never closes that cycle.

hashCode() and equals(Object) sites are now classified synchronous whatever
their implementations are (JavascriptSuspensionAnalysis.isDrivenSignature),
and a generator implementation is run to completion by cn1_ivsDrive, which
now steps through time-slice yields. A body that really blocks there fails
with a named error after withdrawing its monitor entrant. A devirtualized site
whose target is a generator but whose site is synchronous is emitted as _dw.
Kill switch: -Dparparvm.js.drivensigs.off.

A local whose every store is `new X(...)` now has exact type X, so a virtual
call on it goes straight to X's body even when the declared type has
overrides elsewhere (new HashMap() beside LinkedHashMap).
Kill switch: -Dparparvm.js.exactlocals.off.

Gallery, headless Chrome, ?benchCompute=1: hashMapChurn 1488 -> 440ms;
other workloads unchanged. Executable code 12,705,943 -> 12,691,168 bytes.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…s, the blocking accessibility publish

Measured from the Linux leg of the Flutter benchmark (perf profile of the
transpiled gallery's launch) and fixed at the port, so every app gains:

- GDK_GL=disable before gtk_init, unless the environment sets GDK_GL. Opening
  the display otherwise probes GLX for GL-capable visuals (epoxy_glx_version,
  glXQueryServerString, a dlopen of the Mesa driver): ~9% of the launch. The
  port draws with cairo and its 3D backend makes its own EGL context; WebKitGTK
  still renders with it off (checked with WebKitGTK 2.50 under Xvfb). gtk_init
  alone went from 290-630 ms to 90-140 ms in the emulated container.
- Font metrics cached per PangoFontDescription. pango_context_get_metrics
  shapes a sample string for character widths the port never reads, and a
  theme asks for the same descriptions over and over (~7.7% of the launch).
  The cache stores Pango's own first answer, so nothing moves; it is cleared
  when a TrueType font is registered, since a family may then resolve anew.
- The accessibility tree is applied on the GTK thread after the frame, at low
  priority, instead of while the event dispatch thread waits. It builds a GTK
  widget per node (~5% of the launch, all before the first frame). Each
  publication replaces the whole tree, so one still pending when a newer
  arrives is freed unapplied.

A/B on the same translated gallery build (x86_64 under emulation, Xvfb,
1280x720, 10 interleaved launches, load average 53-153): first frame median
860 -> 707 ms, best 846 -> 688 ms. Captures of both, 4 s after the first
frame, are pixel-identical.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…tive overrides

Three costs found profiling the transpiled Flutter gallery in headless Chrome,
plus the test the synchronous hashCode/equals sites were missing.

- A direct call to a synchronous body is now written `impl(_nn(t), ...)` when
  every argument is a local or a literal. Through _dn* the body was a parameter
  of a helper every direct call shares, so its inner call was megamorphic, the
  body was never inlined and the receiver escaped into it: on valueEscape's
  shape (allocate, read two fields through getters) 107ms against 7ms in node.
  The null check now runs before the arguments, which is unobservable only when
  they are pure, so impure argument lists keep _dn*.
- LCONST_0/LCONST_1 emit the literals 0 and 1 rather than _L0/_L1. A long is a
  plain number in the safe range, and a top-level const read is a context load
  V8 did not fold: an accumulator seeded from it kept every _Ladd's type tests,
  2.5x slower on arraySequential's reduction in node.
- installNativeBindings' overrideMethodMaps scanned every class for every
  native, concatenating a prefix per class to find the owner and probing each
  methods map for the key. On the gallery that was 9% of the worker's CPU
  profile at start-up and deoptimized the function 808 times. The owner is now
  found at the name's underscore boundaries and the holders of a key come from
  an index built once.
- JsDrivenHashContentionApp: a synchronized hashCode() whose monitor another
  green thread holds, reached through HashMap.put. It must throw the driver's
  "sync virtual dispatch reached a blocking method" error, never proceed
  unlocked, and leave the monitor and map usable. Verified to fail in all five
  compiler configurations when the driver returns instead of throwing.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Scaled drawing. LinuxImplementation never overrode
isScaledImageDrawingSupported, although the port has always had
drawImageScaled, so Graphics.drawImage(img, x, y, w, h) scaled first: for an
EncodedImage that is EncodedImage.scaled -- decode, scale, re-encode to PNG
(lossy JPEG for an opaque picture), decode again -- on every paint, with
nothing keeping the result. The transpiled Flutter gallery ran 13 PNG encodes
before its first frame (5 cards, 6 icons, the logo twice), all from
PaintSurface.paintDirty. With the override: 0. A capture 4 s after the first
frame differs from the old one in 743 pixels by at most 3 levels -- the JPEG
round trip that no longer happens.

Text widths. cn1MeasureUtf8 now caches (font, exact UTF-8 bytes) -> width,
bounded at 8192 entries (cleared when full) and cleared with the metrics
cache when a TrueType font is registered. 84% of lookups before the gallery's
first frame hit (1508 of 1800), and captures are pixel-identical.

Neither moved the median first frame measurably in the emulated container
(x86_64 under emulation, Xvfb, 14 interleaved launches: 911 / 900 / 917 ms,
spread +-100 ms), where the launch is dominated by other costs.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The same gap 13cbaf7a8f closed on Linux: WindowsImplementation never
overrode isScaledImageDrawingSupported, although the port has always had
drawImageScaled, so Graphics.drawImage(img, x, y, w, h) pre-scaled the
picture -- for an EncodedImage a decode, scale, re-encode to PNG (lossy JPEG
when opaque) and decode on every paint.

drawImageScaled uses the same cn1WinEnsureBitmap, clip push and linear
interpolation as the unscaled drawImage, only into a w x h destination, so
mutable and encoded images and alpha behave exactly as they do unscaled. The
native has no guard for a degenerate rectangle, so the Java override returns
early for w or h <= 0.

Not built here (no Windows toolchain): the Java compiles with JDK 8, the
native signature matches its declaration, and the C-linkage gate passes. CI
verifies. Windows screenshot goldens that draw scaled encoded images may
shift by a few levels, since the lossy re-encode no longer happens -- on
Linux the same change moved 743 pixels by at most 3.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The transpiled Flutter gallery's arraySequential ran 412ms in headless Chrome
against ~180 for the same code in a small application. A natives-syntax probe
in the gallery showed why: its 8M-element int[] was HOLEY_ELEMENTS. V8 records
on each allocation site the most general elements kind an array it made ever
reached and starts later arrays there, and jvm.newArray had ONE `new Array`
site for every array in the program -- so after the first Object[] every int[]
held a boxed HeapNumber per element outside the small-integer range, and every
store allocated. Each element family (int; byte/char/short; boolean; float and
double; long; references) now has its own site.

Array loads and stores inside a loop are emitted as a bounds test statement
followed by a plain element access, `if(a==null||i>>>0>=a.length)_A(a,i);`,
instead of a call to the one _A/_T helper every access in the program shares,
whose element-access feedback is megamorphic in a real application. The
failing case still goes through _A/_T, so NullPointerException and
ArrayIndexOutOfBoundsException are exactly what they were; the value of a
store is still evaluated once, after array and index. Only pure array and
index operands qualify (they are re-read), and only sites between a backward
branch and its target, since the form costs ~30 bytes: +30KB on the gallery.
A conditional expression was measured and rejected -- merging the unboxed
element with _A's tagged result boxed every load (212ms against 38ms for the
statement form in a Chrome micro-benchmark, 98ms for the shared helper).

The non-deferred direct-call path now also writes `impl(_nn(t), ...)`.

Gallery, headless Chrome, 3 interleaved rounds at load 8-12 (ms, origin ->
this): arraySequential 412 -> 206, arrayRandom 835 -> 749, quicksortBench
341 -> 322, stringBuilding 285 -> 275; valueEscape unchanged at 124; the rest
within noise. Checksums identical. Executable code 12,704,437 bytes.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
_L("text") looks the text up in the runtime's literal table on every read, a
dictionary lookup over every literal in the program -- 4,675 distinct ones in
the transpiled Flutter gallery -- repeated per iteration wherever a loop
builds strings. A literal read between a backward branch and its target is now
`(_Lc[n]||(_Lc[n]=_L("text")))`, one slot per distinct text; elsewhere the
lookup runs once and the slot would not pay for its bytes (+11.7KB on the
gallery for the loop sites alone).

Gallery, headless Chrome, stringBuilding, 6 interleaved rounds at load 5-8:
min 272 -> 261ms, median 272 -> 266ms. Small, and inside what this host can
resolve on one workload; kept because it cannot cost anything else.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
8c3d4c1 draws scaled images through the port's drawImageScaled native
instead of decode, scale, re-encode and decode. The old goldens carried the
re-encode path's nearest-neighbour stair steps on upscaled edges; every other
port (Linux, macOS, JavaScript) draws those edges smooth, and so does Windows
now. Native and cross-compiled captures are pixel-identical.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
074cf54 draws scaled images through drawImageScaled instead of decode,
scale, re-encode and decode; only the edges of scaled images move. The x64
and arm64 captures are pixel-identical.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The Linux arm64 gate's hello row crashed (exit 139, round r00) on
737e3b7 and left nothing to autopsy: its stdout and stderr logs --
which carry the ParparVM fault handler's backtrace -- were written under
vm/selfhost/target/perf-gate, which no workflow uploads. The job summary
said "exited with 139" and named a path on a runner that no longer
existed.

A failed row now copies its logs beside perf-results.json (the uploaded
raw/perf directory) and names the copy in the failure reason.

Not reproduced locally: the cull-on self-hosted translator (-O3, thin
LTO, clang 19, arm64 Linux) ran a larger clean translation 120 times and
the cull-off build 90 times without a crash, and a CN1_GC_VERIFY build
reported 0 violations over 207M references. The 56 methods the
reachability cull stubs in the translator binary are called only from
interface thunks that nothing calls; the inline threshold does not reach
the self-hosted binary (compile-dist.sh builds it directly, not through
the generated CMake).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The Windows x64 perf gate measured objectAllocation +33% time and +23%
peak memory after the inline threshold went target-wide. The runtime's
hand-written C -- allocator, collector, write barriers, nativeMethods.c
-- was compiled with it too, and it is five files of a few thousand: not
where the size is, and where a lost inline costs on every allocation.

The generated CMake now reads cn1-source-manifest.txt (written beside
it, after every source exists) at configure time and puts the threshold
in the COMPILE_OPTIONS of the sources the manifest lists as generated.
No file-name guessing: java_io_File_runtime.c is hand-written and named
like a class. List form rather than SHELL:, which per-source options do
not accept.

Checked on the generated Bench project with clang: 147 of 152 compile
commands carry the flag -- exactly the generated set -- and none of
cn1_globals.c, nativeMethods.c, cn1_virtual_thread.c, cn1_win_compat.c,
java_io_File_runtime.c do. CleanTargetIntegrationTest green. The arm64
Linux proxy (clang 19, no LTO) did not reproduce the Windows regression
in any variant (47.5 / 46.2 / 47.3 ms), so the Windows figure has to come
from CI.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The Flutter benchmark's Linux build uses the runner's host cc (no zig
there), which is GCC 13 -- and the clang-only -mllvm threshold never
reached it. CI measured Linux executable code 17.8MB before the size work
and 17.9MB after: no change at all. The reachability cull did run (809
methods + 1,439), but GCC's --gc-sections was already dropping what it
removes; the Mac gained from it because DEAD_CODE_STRIPPING is off there.

GCC's knob is max-inline-insns-single, applied like the clang flag to
the sources the manifest lists as generated. Measured on the transpiled
Flutter gallery's Linux translation:

- object .text (arm64, gcc 12): -O3 26.45MB, single=20 23.90MB.
- linked .text (x86-64, gcc 12, gc-sections, the build as CI runs it):
  18.51MB -> 15.55MB (-16%); Flutter's Linux figure is 16.1MB.
- vm/benchmarks, runtime at plain -O3, 3 interleaved rounds: geomean
  54.91 -> 55.07ms, every workload within noise.

Rejected, each measured: max-inline-insns-auto (shrinks more, but stops
-O3 inlining recursion into itself: recursion 76 -> 116ms), -O2 (same
recursion loss, objectAllocation +9%), turning off cloning, unswitching,
peeling and loop versioning (0.2%).

Checked in the generated build.ninja: all 3,793 generated sources carry
the flag, none of the 28 hand-written runtime and port sources do.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
pipelineFor: called newDefaultLibrary for every pipeline it built, which opens
and parses default.metallib each time -- several times on the first frame --
and, in the non-ARC build, leaked each library object. The cache now keeps the
one library it loaded.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The iOS gallery app was 5.7MB larger than Flutter's before counting a
byte of code, all of it in two places:

- PNGs. Xcode's copypng rewrites every loose PNG into Apple's CgBI form,
  which for photographic images is LARGER: the gallery's 145 PNGs went
  from 14.40MB to 19.98MB. The template now sets COMPRESS_PNG_FILES = NO
  and STRIP_PNG_TEXT = NO on both configurations -- both, because copypng
  treats -strip-PNG-text as "compress" and runs pngcrush -iphone anyway,
  so the first setting alone changed nothing. Nothing here needs CgBI:
  images decode through ImageIO, which reads a standard PNG equally well,
  and Image.isPNG only checks the signature both share.
- WATCHOS_PORT.md and TVOS_PORT.md, which sit beside the port's natives
  and were zipped into nativeios.jar with them, so Xcode put them in the
  resources phase of every iOS app. Excluded at the packaging step (an
  app's own .md resources are untouched).

Gallery, device Release, stripped, as flutter-bench builds it: installed
123.91MB -> 118.24MB (Flutter 115.85MB locally); every asset is now
byte-identical to Flutter's copy.

Cloud builder note (BuildDaemon, separate repo, not changed here): it
builds from the same generated project, so it inherits the template
settings; any IPhoneBuilder path there that rewrites COMPRESS_PNG_FILES or
STRIP_PNG_TEXT, or that archives with an xcconfig setting them, would
undo this.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
ReachabilityCull matched calls by lookup signature only, so x.area() kept every
area() in the program. An instance method is now live only once its class, or a
subclass that inherits it, is allocated. Until then it is held, and it is
released as soon as an allocation happens. Static methods and initializers keep
the signature rule.

Allocation seeds (ReachabilityCull.Allocation):
- NEW in a live method;
- class literals in a live method (NativeLookup.register, Util.register,
  getMenuBarClass...);
- string constants spelling a class's binary name;
- every class named in the hand-written native sources or headers;
- the runtime roots the C side creates (String, the boxes, Thread, the VM's
  exceptions);
- every dependency of a live method holding a fused or custom instruction;
- once Class.newInstance is reachable, every concrete class with a no-arg
  constructor. Class.forName takes names built at run time: NativeLookup's
  "Impl", UIBuilder's resource names, stored geofence/background listener
  names. newInstance can only run a no-arg constructor, so this bounds what
  reflection can create.

A held method that live code names directly (a guarded
((Database) o).close(), for example) keeps its dispatcher. Its body becomes the
culled stub, and its callees stop counting as called
(BytecodeMethod.cullBody, MethodDependencyGraph.removeCalls).
-Dcn1.cullRta=false restores the signature-only answer.

Safety net: with CN1_CULL_TRAP=1 (or -Dcn1.cull.trap=true), every culled stub
aborts and names the method, instead of returning 0. It aborts rather than
throwing an Error, because a catch (Throwable) would swallow an Error. Our
translating workflows set it at workflow level, as CN1_NATIVE_VERIFY is set.
The size and perf benchmarks do not set it, and neither do customer builds.

Measured on the transpiled Flutter gallery, same sources, with and without the
gate:
- iOS (device Release, as flutter-bench builds it): code 21.93 -> 21.23 MB
  (-3.2%); installed 118.24 -> 117.54 MB.
- macOS universal: __TEXT arm64 18.28 -> 17.96 MB, x86_64 19.20 -> 18.85 MB
  (-1.8% each).
- Compute checksums are identical to the baseline.
- The cull removes 2891 methods on iOS (2946 on macOS), against 1031 (1065)
  before; 826 bodies are culled.
The cap is the reflective seed: 1315 of 3866 allocated classes come from it,
because IOSImplementation's background location and fetch paths make
Class.forName reachable in every app.

Also, for clang without LTO (zig cc on Linux), generated sources build at -O2
instead of CMake Release's -O3. Measured with clang at threshold 50: .text
22.43 -> 21.49 MB (-4.2%); vm/benchmarks compute geomean 59.44 -> 57.58 ms;
recursion 161 -> 118 ms. clang-cl already builds /O2, and GCC is untouched,
because -O2 costs recursion ~52% there.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…tance

Until now, the moment Class.newInstance was reachable, every concrete class
with a no-arg constructor counted as allocated. On the transpiled gallery
that was 1315 of 3866 classes. It fired in every app, because the iOS
port's background location, geofence and background-processing paths reach
Class.forName.

newInstance can only act on a Class object the program holds, and there
are three ways to hold one:
- a class literal (already an allocation seed);
- getClass() on an instance (its class made the instance);
- Class.forName, the only one that can name a class nothing allocated.
getSuperclass() yields an ancestor of an allocated class, whose methods the
subtree rule already admits. So the seed now belongs to each live method
that calls forName, and applies once newInstance is live too.

Audit of every forName in the core, the translated ports and JavaAPI
(Mac, Linux, Windows and the flutter/dart runtimes have none). Each site
feeds exactly one type:
- NativeLookup.create -> NativeInterface
- GeofenceManager.getListenerClass -> GeofenceListener
- DeviceRunner.runTest -> UnitTest
- IOSImplementation.Loc.getBackgroundLocationListener -> LocationListener
- IOSImplementation.Loc.getGeofenceListener -> GeofenceListener
- IOSImplementation.runBackgroundProcessing -> BackgroundWorker

A known site seeds only the concrete no-arg classes assignable to its type.
Any other caller (an app's, a cn1lib's) keeps the old answer, every concrete
no-arg class, and the translator prints which method did it. Each call site
now carries a comment pointing at ReachabilityCull.FOR_NAME_SITES.

Every other reflective creation in core was checked and takes a literal or
getClass():
- UIBuilder's component registry;
- Util.register;
- the PropertyBusinessObject types;
- MenuBar via LookAndFeel;
- Facebook/GoogleConnect, AdsService, Socket, CommonProgressAnimations and
  InstantUI;
- DynamicImage and PropertyIndex, which use getClass().

Tests (vm/tests ForNameSeedIntegrationTest):
- For each site, a stand-in with the real class and method name is the
  only live forName, and a listener is created only by a name built at run
  time. The program is built with CN1_CULL_TRAP and run: the listener runs,
  and the other seed classes' methods are culled stubs.
- An unlisted site falls back to every no-arg class and is reported.
- The real core and iOS port are translated with all six sites live, and
  no forName may be reported unnarrowed.
Probes: dropping the BackgroundWorker entry fails the per-site test and the
real-core test. Ignoring the site's type fails the "others are culled"
assertions.

Measured on the gallery, same sources, headless builds:
- iOS code 21.23 -> 20.42 MB (Flutter 20.00). Before type-aware reachability
  it was 21.93. Installed 117.54 -> 116.73 MB.
- macOS __TEXT arm64 17.96 -> 17.83 MB, x86_64 18.85 -> 18.72 MB.
- Allocated classes 3866 -> 3545 on iOS. The cull removes 4323 methods
  (+919), against 2891 (+771).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The window was queued onto the main queue, so it could not be built until
[NSApp run] had finished launching -- on the GitHub macOS runner ~264 ms into
the process -- and the first frame waited behind it. The generated main now
calls CN1MacBuildMainWindowBeforeRun right after dispatching the VM boot, so
the build overlaps the boot instead of delaying it (the reason the old comment
gave for not building inline). Weakly linked: a port without the function keeps
the queued build.

flutter-bench macOS on CI (diagnostic run 36613510954): window built at
~130 ms instead of ~264, the Java first paint 125 ms -> 7-32 ms, cold start
344 ms vs Flutter 382 ms (1.11x; was 0.87x).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…cons

Measured on the transpiled Flutter gallery, same sources, headless device builds
(iOS: Release, iphoneos, as flutter-bench builds it):

| | before | after | Flutter |
|---|---|---|---|
| iOS code | 20.42 MB | 19.37 MB | 20.00 MB |
| iOS installed | 116.73 MB | 115.03 MB | 115.85 MB |
| macOS __TEXT arm64 | 17.83 MB | 16.70 MB | |
| macOS __TEXT x86_64 | 18.72 MB | 17.52 MB | |

The reachability cull removes 5614 methods (+1290 after the class cull) on
iOS, against 4323 (+919). The feature packages the gallery never touches are
gone from the binary: home, call, nearby and health.

Four narrowings in ReachabilityCull, active with the type-aware cull. Each
has its own switch, and each one, switched off, fails
NativeScopeIntegrationTest:

- NativeBodies (-Dcn1.nativeBodyScope=false). A name inside the C body of a
  port's Java native, or inside a static C helper, is a call that body makes.
  It is followed once the native is live, and is no longer a root. The
  #else trampolines of every switched-off feature call their callbacks to
  report NOT_SUPPORTED, so every app kept every feature. Static helpers
  only, because their callers are all in the same file. Functions with
  external linkage stay roots, as do the VM runtime's own natives.
  #include/#import lines no longer count as naming a class.
- NativeFeatureFilter (-Dcn1.nativeFeatureFilter=false). A branch that a
  CN1_INCLUDE_ switch compiles out names nothing. So does a CN1_ macro whose
  every #define sits in such a branch, for example CN1_VPN_HAS_NE. Only OFF
  is ever concluded; a condition it cannot evaluate keeps all its branches.
- Conditional <clinit> (-Dcn1.cullClinit=false). A static initializer runs
  only once its class is touched: its code runs, its statics are read or
  written, or it is allocated. The IOS*Callbacks dead-code guards call every
  callback from <clinit>.
- StaticCalls (-Dcn1.cullStaticOwner=false). A static call reaches only the
  class it names or that class's superclasses, not every static with the
  same name and descriptor. A name reached some other way (a fused or
  rewritten instruction, invokedynamic) keeps the signature rule.

A method that compiled C still names but that can never run stays declared, as
the culled stub, so the natives still link. With CN1_CULL_TRAP that stub
aborts.

The builder now puts the app icon only in the AppIcon asset catalog. Loose
copies of every size, Icon.png and iTunesArtwork (0.6 MB) went into the
bundle root for the top-level CFBundleIconFiles of iOS 6 and older. No
supported deployment target reads those files; they take the icon from
CFBundleIcons/CFBundleIconName, which actool writes. The key is dropped from
the Info.plist template, and IPhoneBuilderIconsTest pins it. The BuildDaemon
IPhoneBuilder still copies them and needs the same change.

The ZooZ strings bundle moves from nativeSources to the Zooz SDK directory.
Only a build that links ZooZ uses it, and only BuildDaemon ever enables ZooZ,
by unzipping zoozIosSources.jar, which now carries the bundle.

Verified with CN1_CULL_TRAP on:
- translator tests: 203 run, 0 failures, 11 skipped;
- the torture gauntlet is green;
- the Linux gallery, built in the amd64 container and run under Xvfb there,
  went through 17 scripted click paths with no culled method called, and its
  compute checksums match.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Code grew 4,720,336 -> 4,772,660 bytes (+1.1%) with features added since the
baseline was recorded; still 4.5x under Flutter's 20.6 MB. Only code_bytes
moves: the start-up and memory rows stay at the multi-run envelope they were
tuned to.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
ThemeMode.CUSTOM ("the theme the application ships, rather than a
platform one") is now honoured everywhere, generically -- nothing here
knows about Flutter:

- iOS / macOS: IOSImplementation.installNativeTheme() returns before the
  chain and hasNativeTheme() answers false for custom. Custom used to fall
  to the bottom of the chain and install the pre-flat iPhoneTheme.res.
- Android: the mode resolution moves into nativeThemeMode(), which also
  recognizes custom from and.themeMode / cn1.androidTheme / nativeTheme;
  hasNativeTheme() answers false and installNativeTheme() does nothing.
  Custom used to fall to androidTheme.res.
- Builders: NativeThemes.themeFor answers null and androidThemesFor an
  empty set for custom, so no theme .res is packaged. IPhoneBuilder maps
  nativeTheme=custom to ios mode custom (it became auto, iOS 7), and the
  macOS whitelist accepts custom (it became modern).
- Linux / Windows native: those ports install their bundled theme from
  init() whenever the file exists, so the builders leave
  linuxNativeTheme.res / windowsNativeTheme.res out for custom.
- JavaSE: ios/and.themeMode=custom resolve to no simulator theme, and
  nativeTheme=custom reaches the desktop as "no platform theme" (the
  packaged app and the desktop wrapper default), the way native does.
- Hint docs: custom added to the ios/and.themeMode value patterns and the
  ThemeMode / Build / Mac / DesktopBuild wording.

flutter-bench: the generated app sets includeNativeBool: false in
theme.css, and every build passes nativeTheme / ios.themeMode /
and.themeMode = custom (the archetype's settings set the platform ones
to modern, and they outrank nativeTheme). The porting guide describes
both settings as the generic option they are.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…first use

Two costs every Flutter-runtime launch paid before its first frame:

1. The Linux and Windows ports installed their native theme eagerly in
   init(). FlutterUI then installs CN1FlutterMaterialTheme.res with
   setThemeProps, which replaces the whole table (the theme sets
   @includeNativeBool=false), so that install was thrown away. It is now
   registered with the new UIManager.setDeferredBaseTheme and runs only if
   something reads the theme before any other theme is installed. An app
   with no theme of its own still gets it. setThemeProps drops it, and
   addThemeProps layers on top of it.

2. DefaultLookAndFeel.refreshTheme ends every theme install on every
   platform, and it built the check box, radio button and combo arrow
   glyphs each time: it loaded the Material icon font and rendered eight
   FontImages. These are now built the first time a getter, a paint or a
   preferred-size call asks for them, using the same code and the same
   theme. A setter called after a refresh wins over the late build.

Measured on the Flutter gallery, ParparVM-translated, Linux container
(x86 under emulation), with 8 interleaved launches per arm and medians:
  A HEAD:        native install at init 59.7 ms + Material install 22.3 ms
  B deferred:    Material install 55.8 ms (it is now the first install)
  C B + glyphs:  Material install 45.8 ms
  Total theme work before the first frame: 82 -> 46 ms.
  Cold start to BENCH:FIRSTFRAME (14 interleaved runs, median):
  483 / 475 / 468 ms.
  Captures 4 s after the first frame: A == B == C, 0 pixels differ.
Headless JVM (JDK 8, 10 fresh JVMs per arm), setThemeProps of
CN1FlutterMaterialTheme.res:
  first install 13.4 -> 9.9 ms, warm 1.13 -> 0.73 ms.
On macOS and iOS the Material install is already the first install, so
only the glyph half applies there. It is the same share of refreshTheme.

Tests: UIManagerDeferredBaseThemeTest and DefaultLookAndFeelLazyGlyphsTest.
Full core-unittests: 7482 run, 0 failures.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant