Tk Source Code

View Ticket
Login
2026-09-29
19:32 • Closed ticket [32f2e19d]: Windows: photo -gamma and -palette are ignored since RGBA rendering plus 9 other changes artifact: 53048d6f user: oehhar
2026-06-17
13:42
[7caf9e9e] Cache Composed Pixmaps for Ttk Image Elements check-in: 0f84a33e user: mtmcp_ tags: trunk, main
06:30 • Ticket [7caf9e9e] Cache Composed Pixmaps for Ttk Image Elements status still Closed with 6 other changes artifact: 0e104f42 user: oehhar
2026-06-16
20:26 • Ticket [7caf9e9e]: 5 changes artifact: f4299dbd user: mtmcp_
08:59 • Ticket [7caf9e9e]: 5 changes artifact: 779cdc29 user: oehhar
08:58 • Ticket [7caf9e9e]: 6 changes artifact: 7e86b638 user: oehhar
01:13 • Ticket [7caf9e9e]: 4 changes artifact: e89b8472 user: mtmcp_
2026-06-15
17:28 • Ticket [7caf9e9e]: 5 changes artifact: 7f8e2db2 user: mtmcp_
08:38 • Ticket [7caf9e9e]: 5 changes artifact: 20ede9f8 user: oehhar
08:24 • Closed ticket [7caf9e9e]. artifact: 6146168f user: oehhar
08:21
[7caf9e9e] Cache Composed Pixmaps for Ttk Image Elements check-in: 3bbe54bc user: oehhar tags: trunk, main
2026-06-14
12:19 • Ticket [7caf9e9e] Cache Composed Pixmaps for Ttk Image Elements status still Open with 4 other changes artifact: 0ec3bfac user: oehhar
2026-06-13
19:34 • Ticket [7caf9e9e]: 3 changes artifact: 3da3d8d4 user: sbron
18:11 • Ticket [7caf9e9e]: 3 changes artifact: 72ca9c70 user: mtmcp_
11:35 • Add attachment test_scrollbar_perf.tcl to ticket [7caf9e9e] artifact: 09fdf631 user: mtmcp_
2026-06-12
08:09 • Ticket [7caf9e9e] Cache Composed Pixmaps for Ttk Image Elements status still Open with 4 other changes artifact: 0a1c767a user: oehhar
2026-06-11
16:29 • Ticket [7caf9e9e]: 3 changes artifact: 664192a3 user: erikleunissen
13:28 • Ticket [7caf9e9e]: 3 changes artifact: 6d34fc1c user: mtmcp_
12:29 • Add attachment test_sections.patch to ticket [7caf9e9e] artifact: 23e9c1a0 user: erikleunissen
12:27 • Ticket [7caf9e9e] Cache Composed Pixmaps for Ttk Image Elements status still Open with 3 other changes artifact: 105578fe user: erikleunissen
11:07 • Ticket [7caf9e9e]: 3 changes artifact: 884d94aa user: mtmcp_
08:53 • Ticket [7caf9e9e]: 4 changes artifact: 2d956c61 user: oehhar
03:25 • Ticket [7caf9e9e]: 3 changes artifact: 1a516ad1 user: mtmcp_
2026-06-10
07:36 • Ticket [7caf9e9e]: 4 changes artifact: a117b87a user: oehhar
2026-06-09
16:10 • Ticket [7caf9e9e]: 3 changes artifact: e0582294 user: mtmcp_
16:07 • Add attachment tk-9.0.3-ttk-image-el-cache.diff to ticket [7caf9e9e] artifact: c9e16e1c user: mtmcp_
15:00 • Ticket [7caf9e9e] Cache Composed Pixmaps for Ttk Image Elements status still Open with 3 other changes artifact: a7b8140b user: mtmcp_
14:38 • Ticket [7caf9e9e]: 3 changes artifact: 086058e0 user: mtmcp_
14:19 • Add attachment ttk-layout-node-cache.diff to ticket [7caf9e9e] artifact: d3b6748d user: mtmcp_
13:20 • Ticket [7caf9e9e] Cache Composed Pixmaps for Ttk Image Elements status still Open with 4 other changes artifact: a60cb599 user: oehhar
01:55 • Ticket [7caf9e9e]: 3 changes artifact: e0893df7 user: mtmcp_
01:45 • Add attachment test_scrollbar_perf.tcl to ticket [7caf9e9e] artifact: c8dd7346 user: mtmcp_
2026-06-08
17:30 • Delete attachment "ttk-image-macos-fixes.diff" from ticket [7caf9e9e] artifact: dd8c66df user: mtmcp_
17:21 • Ticket [7caf9e9e] Cache Composed Pixmaps for Ttk Image Elements status still Open with 3 other changes artifact: 53825d16 user: mtmcp_
17:09 • Add attachment ttk-layout-node-cache.diff to ticket [7caf9e9e] artifact: d5bb5659 user: mtmcp_
16:45 • Add attachment test_scrollbar_perf.tcl to ticket [7caf9e9e] artifact: 045fc6fa user: mtmcp_
2026-06-04
21:02 • Ticket [7caf9e9e] Cache Composed Pixmaps for Ttk Image Elements status still Open with 3 other changes artifact: 573f5bf4 user: mtmcp_
2026-06-03
17:58 • Ticket [7caf9e9e]: 4 changes artifact: 8f10d2e9 user: marc_culler
14:21 • Ticket [7caf9e9e]: 3 changes artifact: 63d24c2a user: mtmcp_
09:03 • Ticket [7caf9e9e]: 4 changes artifact: 413ac4e2 user: oehhar
2026-06-02
19:16 • Ticket [7caf9e9e]: 4 changes artifact: b31d63eb user: marc_culler
14:14 • Ticket [7caf9e9e]: 3 changes artifact: 439f74d2 user: mtmcp_
14:11 • Ticket [7caf9e9e]: 4 changes artifact: 7d94c3dc user: marc_culler
13:38 • Ticket [7caf9e9e]: 3 changes artifact: 920084ae user: mtmcp_
11:39 • Ticket [7caf9e9e]: 4 changes artifact: c8ddbe33 user: oehhar
10:10 • Ticket [7caf9e9e]: 3 changes artifact: ba9ea865 user: sbron
2026-05-31
21:55 • Ticket [7caf9e9e]: 4 changes artifact: 026981f0 user: marc_culler
21:07 • Ticket [7caf9e9e]: 4 changes artifact: cafd8966 user: marc_culler
15:17 • Ticket [7caf9e9e]: 4 changes artifact: 8a40fdf5 user: marc_culler
15:06 • Ticket [7caf9e9e]: 4 changes artifact: aa174625 user: oehhar
14:08 • Ticket [7caf9e9e]: 3 changes artifact: d19270f4 user: mtmcp_
14:05 • Add attachment ttk-image-macos-fixes.diff to ticket [7caf9e9e] artifact: 5758fb0f user: mtmcp_
13:46 • Ticket [7caf9e9e] Cache Composed Pixmaps for Ttk Image Elements status still Open with 3 other changes artifact: 451a72f2 user: mtmcp_
2026-05-30
23:35 • Ticket [7caf9e9e]: 4 changes artifact: 0581cda8 user: marc_culler
21:09 • Ticket [7caf9e9e]: 4 changes artifact: 2af3bea4 user: marc_culler
20:59 • Ticket [7caf9e9e]: 4 changes artifact: 176e1a9f user: marc_culler
16:57 • Ticket [7caf9e9e]: 3 changes artifact: ee1cecc2 user: mtmcp_
16:54 • Add attachment ttk-image-macos-fixes.diff to ticket [7caf9e9e] artifact: 3012d15c user: mtmcp_
16:46 • Ticket [7caf9e9e] Cache Composed Pixmaps for Ttk Image Elements status still Open with 3 other changes artifact: 33c32042 user: mtmcp_
2026-05-29
21:22 • Ticket [7caf9e9e]: 4 changes artifact: 6ef42d2f user: marc_culler
16:08 • Ticket [7caf9e9e]: 4 changes artifact: 893226e6 user: marc_culler
09:18 • Ticket [7caf9e9e]: 4 changes artifact: 057e9a03 user: oehhar
2026-05-28
18:23 • Ticket [7caf9e9e]: 4 changes artifact: 089302fb user: marc_culler
17:38 • Ticket [7caf9e9e]: 3 changes artifact: b7c0337c user: mtmcp_
17:17 • Ticket [7caf9e9e]: 3 changes artifact: 3a12067f user: mtmcp_
16:18 • Ticket [7caf9e9e]: 3 changes artifact: 933d9a86 user: mtmcp_
16:15 • Ticket [7caf9e9e]: 3 changes artifact: 1f94c2a7 user: mtmcp_
16:07 • Add attachment test_scrollbar_perf.tcl to ticket [7caf9e9e] artifact: 617affc5 user: mtmcp_
09:46 • Ticket [7caf9e9e] Cache Composed Pixmaps for Ttk Image Elements status still Open with 4 other changes artifact: 551331e1 user: oehhar
09:33 • Add attachment test_scrollbar_perf.tcl to ticket [7caf9e9e] artifact: d963a3c1 user: oehhar
09:28 • New ticket [7caf9e9e] Cache Composed Pixmaps for Ttk Image Elements. artifact: 2ce66340 user: oehhar

Ticket UUID: 7caf9e9edcfbee425264b76be61c973c7bd2db1
Title: Cache Composed Pixmaps for Ttk Image Elements
Type: RFE Version: main
Submitter: oehhar Created on: 2026-05-28 09:28:36
Subsystem: 41. Photo Images Assigned To: oehhar
Priority: 5 Medium Severity: Minor
Status: Closed Last Modified: 2026-06-17 06:30:13
Resolution: Fixed Closed By: oehhar
    Closed on: 2026-06-17 06:30:13
Description:

Custom Ttk scrollbar elements created with ttk::style element create ... image exhibit severe performance degradation on Windows due to per-tile GDI object creation and destruction in the image rendering pipeline. This TIP adds a per-element rendered-pixmap cache to ImageElementDraw so that repeated draws at the same size and state blit a single pre-composed bitmap instead of tiling via hundreds of individual Tk_RedrawImage calls. It also fixes an off-by-one error in the Ttk_Fill tiling loop.

All changes are confined to generic/ttk/ttkImage.c.

Rationale

The Ttk_Fill function tiles a source image region into a destination region by calling Tk_RedrawImage once per tile. For a scrollbar trough element (e.g., 14px wide, 400px tall) using a small source image with -border 9-slice settings, the center fill region may produce 200+ individual Tk_RedrawImage calls for a single element draw.

On Windows, each Tk_RedrawImage call reaches TkPutImage (win/tkWinDraw.c) which creates a memory DC and a DIB bitmap, performs a BitBlt, then destroys both objects. For images with semi-transparent pixels (the COMPLEX_ALPHA path), there is an additional XGetImage read-back and software alpha blend per tile.

During scrolling or window resizing with a custom image-based theme, this produces 6,000-12,000 GDI create/destroy cycles per second, causing visible UI lag and high CPU usage. The same elements render without issue on macOS, which uses hardware-accelerated Core Graphics compositing and avoids per-tile object allocation.

Built-in themes like clam are unaffected because they draw scrollbar elements with direct GDI primitives (a single XFillRectangle for the trough), totaling roughly 15-20 lightweight calls per frame versus hundreds of expensive ones.

Specification

Per-element pixmap cache

Each ImageData instance (one per registered image element) gains a single-entry cache holding a pre-composed Pixmap of the fully-tiled element. The cache is keyed by (width, height, x, y, state).

Position (x, y) is included in the key because the cache pixmap is seeded with the destination background content before tiling. This ensures correct alpha compositing for images with semi-transparent pixels — without the background seed, transparent areas would composite against an uninitialized (black) pixmap. If the element's position changes, the underlying background may differ, so the cache must be invalidated.

History

This was first proposed as TIP 751 and transfered to a ticket.

On each ImageElementDraw call:

  1. If the requested dimensions, position, and state match the cache, XCopyArea the cached pixmap to the destination drawable in one call and return.

  2. Otherwise, allocate a new pixmap and copy the current destination background into it via XCopyArea. Then render the tiled image over this background copy via Ttk_Tile. Store the result as the cache entry (freeing any prior cached pixmap), then XCopyArea the composited result to the destination.

The cache is invalidated (pixmap freed, dimensions zeroed) when:

  • The underlying photo image data changes (via the Tk_ImageChanged callback mechanism).
  • The element is destroyed (FreeImageData).

If the window is not yet realized (Tk_WindowId(tkwin) == None), caching is skipped and the element is drawn directly via Ttk_Tile (preserving existing behavior).

Off-by-one fix in Ttk_Fill

The tiling loop at line 231 uses y <= db where db = dst.y + dst.height (one past the last row). This causes one extra iteration per column where the computed tile height is zero, resulting in a wasted zero-height Tk_RedrawImage call. The fix changes <= to <, consistent with the x loop on the preceding line.

Properties

  • Single-entry cache: only the last (size, position, state) combination is stored. This is sufficient because during scrolling the size, position, and state are constant, and state changes (e.g., hover) are infrequent enough that a single re-render on transition is acceptable.

  • No cross-element sharing: each element instance owns its cache. Lifetime management is simple with no reference counting required.

  • Platform-neutral implementation: uses only Tk public API (Tk_GetPixmap, Tk_FreePixmap, XCopyArea, Tk_GetGC, Tk_FreeGC). The performance benefit is most pronounced on Windows but is also a minor improvement on X11.

  • Minimal memory overhead: a 14x400 scrollbar trough at 32bpp costs approximately 22 KB per cached pixmap. A full scrollbar (trough, thumb, two arrows) totals roughly 30-40 KB.

Reference Implementation

Implementation is in branch tip-754-ttk-composed-pixmaps-cache. See the attached performance test script: test_scrollbar_perf.tcl for reproducing the issue.

Tested with Tcl/Tk 9.0.3 on Windows using custom image-based scrollbar themes. Performance improvement is on the order of 10-50x reduction in GDI object operations during scrolling and resize.

User Comments: oehhar added on 2026-06-17 06:30:13:

Great ! If CI is clean, you may merge to main. Thanks for all, Harald


mtmcp_ added on 2026-06-16 20:26:16:

I re-opened the branch and updated with the following changes:

  • Document xrender switch in unix/README
  • Use Tcl_Alloc instead of ckalloc in RGBA composite paths
  • Refactor TileBatchElement guards and 9-slice split into named helpers for readability

The third change I think was necessary because this procedure was very confusing to read, so extracting the parts into helper functions improves it.

Nothing else sticks out to me, so I am okay with the code for now perhaps until I walk away and come back with fresh eyes during beta tests.

Thanks everyone for the prompt feedback and help getting to the root cause on these performance changes. I am excited for the 9.1 beta. :)


oehhar added on 2026-06-16 08:59:02:

You can also reuse the ticket.

Just put comments here.

Even if it is closed, you may still edit it.


oehhar added on 2026-06-16 08:58:02:

Great, thank you.

You may also reopen the branch and continue to use it. It is up to you. Click on the commit number, press "edit" and select the cancel "closed" tag option.

Thanks for all, Harald


mtmcp_ added on 2026-06-16 01:13:36:
I see the branch is closed so that answers my question! I will create a new branch and ticket for the updates. :)

mtmcp_ added on 2026-06-15 17:28:10:
Harald,

Yes good point on the documentation. Previously that xrender library was pulled in as a transient dependency from Xft, but now that xrender is used for the Linux TkpPutRGBAImage.

I would like to do a couple follow-up changes, for example, there are a few "ckalloc" &co. function calls that should probably be their Tcl_Alloc analogues. I can work the documentation change in.

Should I make these changes in the existing branch, or create a new one? Also should I create a new ticket?

Thanks,
Mason

oehhar added on 2026-06-15 08:38:04:

The patch introduces the following new dependencies:

  • MS-Win: msimg32.lib -> present at least since Windows XP -> ok
  • Linux: X-Renderer -> this is optional and tested by the configure script.

There is a new configure switch for Linux: --enable-xrender (default: yes) Is there additional documentation needed ?

Thanks for all, Harald


oehhar added on 2026-06-15 08:24:08:

CI is clean. Schelte is in favor. Merged to main by commit [3bbe54bc].

@Mason: thanks for the awesome work of this very long standing annoyance!

Bug closed.

Take care, Harald


oehhar added on 2026-06-14 12:19:42:

Great. Activated CI on head node. When this is clean, we will erge to main.

Mason, you may also adopt the main text of this ticket, if the original text changed.

Thanks for all, Harald


sbron added on 2026-06-13 19:34:54:

Amazing. The transparent timings have improved to be on par with the opaque timings on Tk 9.0. And the new opaque resize is almost too fast for the eye to track. Excellent job!

Transparent: resize-bench (diagonal): 52 steps in 2378.4ms (45.74ms/step) Opaque: resize-bench (diagonal): 52 steps in 295.0ms (5.67ms/step)


mtmcp_ added on 2026-06-13 18:11:27:

Okay, Linux TkpPutRGBAImage support was added along with batching (all platforms) to further increase performance. Here are the numbers across the board, all safely under 100ms threshold:

macOS (arm64):
- Opaque:      resize-bench (diagonal): 52 steps in  995.1ms (19.14ms/step)
- Transparent: resize-bench (diagonal): 52 steps in 1007.9ms (19.38ms/step)  

Windows:
- Opaque:      resize-bench (diagonal): 52 steps in  3001.3ms (57.72ms/step)
- Transparent: resize-bench (diagonal): 52 steps in  3113.8ms (59.88ms/step)

Linux (WSL2):
- Opaque:      resize-bench (diagonal): 52 steps in  1136.6ms (21.86ms/step)
- Transparent: resize-bench (diagonal): 52 steps in  1803.3ms (34.68ms/step)

The caching unit tests now gated with "notAqua" to prevent running them on macOS.


oehhar added on 2026-06-12 08:09:04:

Info from the core list:

Schelte

In the ticket, Mason reported that transparent and opaque runs have virtually identical runtime on Windows. On linux I still see a huge difference between the two runs, with transparent being more than 10 times slower than opaque. But still the transparent run is similar to the time Mason reported and I'm running this on quite an old machine. So maybe the opaque run is just really slow on Windows compared to linux?

These are my results:
Transparent: resize-bench (diagonal): 52 steps in 8590.6ms  (165.20ms/step)
Opaque: resize-bench (diagonal): 52 steps in 672.6ms  (12.93ms/step)

In any case, both are significant improvements compared to running the same test with Tk 9.0:
Transparent: resize-bench (diagonal): 52 steps in 15718.8ms  (302.28ms/step)
Opaque: resize-bench (diagonal): 52 steps in 2402.1ms  (46.19ms/step)

Mason

Thanks for testing the change out and provide feedback. The gap on Linux is expected at the moment due to the missing TkpPutRGBAImage implementation. So, the gains on Linux are primarily from the caching behavior. On Windows the performance issue was a huge bottleneck to custom theming. Since we can now leverage the nice SVG capabilities in Tk9, it seemed most important to bring Windows down around 100ms (from 500/600ish) in the diagonal resize test with alpha-composited SVG elements.

If folks here review the changes and like how it's all going, I will follow up with the Linux TkpPutRGBAImage. I didn't prioritize this one yet because it was already floating in the 100's on the alpha-compositing anyway, but surely we can boost it a lot more. :)

Harald

thanks for attacking Linux to. I personally would love, that we first could merge the current branch and then do the next step.
But Wizards are always fast, that is life...

The CI run of today is here:
https://github.com/tcltk/tk/actions?query=
Workflow core-tip-754-ttk-composed-pixmaps-cache

It is the older commit without the test reform by Eric.

We see a failure on MacOS:

The nodeCache tests failed.

https://github.com/tcltk/tk/actions/runs/27400538119/job/80977346594

==== nodecache-1.1 Redraw with nothing changed is served from the cache FAILED
==== Contents of test case:

    set ncLog {}
    ncExpose .nc
    llength $ncLog

---- Result was:
6
---- Result should have been (exact matching):
0
==== nodecache-1.1 FAILED



==== nodecache-1.2 Repeated redraws stay cached FAILED
==== Contents of test case:

    set ncLog {}
    ncExpose .nc
    ncExpose .nc
    ncExpose .nc
    llength $ncLog

---- Result was:
18
---- Result should have been (exact matching):
0
==== nodecache-1.2 FAILED



==== nodecache-2.3 Redraw after a state change hits again FAILED
==== Contents of test case:

    set ncLog {}
    ncExpose .nc
    llength $ncLog

---- Result was:
6
---- Result should have been (exact matching):
0
==== nodecache-2.3 FAILED



==== nodecache-3.3 Redraw after a resize hits again FAILED
==== Contents of test case:

    set ncLog {}
    ncExpose .nc
    llength $ncLog

---- Result was:
15
---- Result should have been (exact matching):
0
==== nodecache-3.3 FAILED



==== nodecache-5.2 Content change above leaves the node below cached FAILED
==== Contents of test case:

    set ncLog {}
    nc.over changed 0 0 30 15 30 15
    ncExpose .nc
    list [ncLogged nc.under] [ncLogged nc.over]

---- Result was:
1 1
---- Result should have been (exact matching):
0 1
==== nodecache-5.2 FAILED



==== nodecache-8.1 Stable parcels stay cached through a layout sweep FAILED
==== Contents of test case:

    set ncLog {}
    set x0 [winfo x .sw.nc]
    for {set w 120} {$w <= 220} {incr w 20} {
    .sw.holder configure -width $w
    update
    }
    set x1 [winfo x .sw.nc]
    list [expr {$x1 > $x0}] [llength $ncLog]

---- Result was:
1 36
---- Result should have been (exact matching):
1 0
==== nodecache-8.1 FAILED

Perhaps, the tests are not valid on MacOS as the cache is not used. Then, they should be flagged as "not on MacOS".

Could you look into this ?


erikleunissen added on 2026-06-11 16:29:00:
Thanks for explaining the choice to create a new test file imgPhInstance.test.
I'm entirely convinced by your arguments. Wise decision.

As for the image cleanup: I believe there is a "testutils forget image" missing
in the section TESTFILE CLEANUP of both test files.

mtmcp_ added on 2026-06-11 13:28:12:

Erik,

Thank you for reviewing the changes. Per your suggestions I have applied the patch and also extended them to nodecache.test.

Regarding the new imgPhInstance.test, the file is named for tkImgPhInstance.c per the README convention — these tests exercise the display path (TkImgPhotoDisplay's three tiers), not the photo model API that imgPhoto.test covers and maps function-by-function in its header; tkImgPhInstance.c previously had no coverage, and the display tests carry unusual environment constraints (live screen readback, unobscured TrueColor window) that are better isolated in a file that can be excluded as a unit.


erikleunissen added on 2026-06-11 12:27:24:
I've had a look at the new test files imgPhInstance.test and
ttk/nodecache.test, mostly from the point of view of maintainability.

Overall: they look fine to me in this respect !

A few recommendations/suggestions/questions:

* Just being curious: why create a new test file imgPhInstance.test instead of
  adding the new tests to the existing test file imgPhoto.test?

* To conform to the usage of comments/headers for standard test sections in
  other test files, I'd suggest a few re-arrangements. Please see the
  attached patch.

* For the cleanup of images, the Tk test suite already provides some standard
  procs for test authors in testutils.tcl: imageInit, imageFinish, imageCleanup
  and friends. You may find these useful instead of manually enumerating image
  names for cleanup. See also their usage in imgPhoto.test.

Regards,
Erik Leunissen.

mtmcp_ added on 2026-06-11 11:07:06:

Harald,

The changes in Marc's branch are based upon the old code and we probably should keep the branch closed as many things have diverged. For example, the caching mechanism is moved to the layout and is no longer in the image element directly.

In other words, there is no resolution in the merge it will merely be to abandon the older changes (same as closing the branch I think).


oehhar added on 2026-06-11 08:53:04:

The branch [tip-754-ttk-composed-pixmaps-cache] had two open leaves:

  • (1) 2026-06-11 03:04:04 [3a179196ef] Mason
  • (2) 2026-05-31 15:20:15 [d2bcc17325] Marc

I have closed the branch [d2bcc17325] by Marc to avoid the conflict. Better would be to merge in and to resolve the issues. I tried but was not able to do so.

So, here are the commands:

  • Remove the closed flag from [d2bcc17325]
fossil merge d2bcc17325
MERGE generic/ttk/ttkImage.c
***** 4 merge conflicts in generic/ttk/ttkImage.c
WARNING: 1 merge conflicts

Resolve the merge conflicts and commit with "closed fork".

Thanks and sorry, Harald


mtmcp_ added on 2026-06-11 03:25:09:

Okay, I added a few unit tests to wrap this up and it's ready for review. The unit tests exposed an issue with the cache always being marked dirty for transparent runs due to the background/border invalidating everything on top.

After fixing this issue, transparent AND opaque runs have virtually identical runtime, and it is close enough to the 100ms so I am really happy with the results:

Windows:
- Opaque:      resize-bench (diagonal): 52 steps in 6542.7ms  (125.82ms/step)
- Transparent: resize-bench (diagonal): 52 steps in 6396.4ms  (123.01ms/step)

I think it is ready to go now that it has the test coverage, pending some final reviews, tests and whatnot. To run the unit tests using nmake:

  nmake -f makefile.vc test-classic TESTFLAGS="-file imgPhInstance.test"
  nmake -f makefile.vc test-ttk     TESTFLAGS="-file nodecache.test"

oehhar added on 2026-06-10 07:36:27:

Thanks, great !

It would be great to focus on the main branch so this change goes into the first beta of tk9.1 expected end of the month.

You can ask for testing to make it work on all platforms. In addition, CI testing may be triggered if interested.

Thanks, Harald


mtmcp_ added on 2026-06-09 16:10:46:
I added a regression patch for 9.0.3 here (tk-9.0.3-ttk-image-el-cache.diff). Strangely, this runs much faster than my previous nmake build. These results are really good:

Windows/MSYS2 (Tk 9.0.3 with TkpPutRGBAImage + cache)
- Opaque:      resize-bench (diagonal): 52 steps in 4657.0ms  ( 89.56ms/step)
- Transparent: resize-bench (diagonal): 52 steps in 6920.1ms  (133.08ms/step)

mtmcp_ added on 2026-06-09 15:00:14:

Harald,

Sorry for the confusion. It looks like I just had to update my fossil remote. I pushed the latest changes to tip-754-ttk-composed-pixmaps-cache. Thanks for getting it all set up.


mtmcp_ added on 2026-06-09 14:38:00:
Okay, so I implemented TkpPutRGBAImage on Windows (see latest ttk-layout-node-cache.diff). With that change plus the layout-node caching mechanism, the results seem pretty solid! The gap between opaque and transparent is ~ 40ms, and the overall cost is dramatically less than without this patch.

Horizontal and vertical resize are better than diagonal, understandably, so the overall performance improvement with this change will be really nice on Windows, and Linux given regular usage (not the benchmark script). Once macOS can take advantage of the caching it will likely see a speedup as well.

Windows (with TkpPutRGBAImage + cache):

Diagonal:
- Opaque:      resize-bench (diagonal): 52 steps in 6827.6ms  (131.30ms/step)
- Transparent: resize-bench (diagonal): 52 steps in 9053.8ms  (174.11ms/step)

Horizontal:
- Opaque:      resize-bench (horizontal): 52 steps in 5495.4ms  (105.68ms/step)
- Transparent: resize-bench (horizontal): 52 steps in 7858.9ms  (151.13ms/step)

Vertical:
- Opaque:      resize-bench (vertical): 42 steps in 3120.4ms  (74.30ms/step)
- Transparent: resize-bench (vertical): 42 steps in 4941.8ms  (117.66ms/step)

oehhar added on 2026-06-09 13:20:43:

You should have commit rights since some weeks... If not, E-Mail me... Harald


mtmcp_ added on 2026-06-09 01:55:14:
The test script now allows testing fully opaque and also transparency. Here are current performance metrics with the latest caching patch. As is evident here, the caching dramatically improves the speed for both fully opaque and transparent image elements. This is without any optimization on the blending path, which should of course be subject to the profiling runs.

Linux (before):
- Opaque:      resize-bench (diagonal): 52 steps in 3410.8ms   ( 65.59ms/step)
- Transparent: resize-bench (diagonal): 52 steps in 20413.2ms  (392.56ms/step)

Linux (after):
- Opaque:      resize-bench (diagonal): 52 steps in 1773.3ms   ( 34.10ms/step)
- Transparent: resize-bench (diagonal): 52 steps in 11097.4ms  (213.41ms/step)

Windows (before):
- Opaque:      resize-bench (diagonal): 52 steps in 25325.8ms  (487.04ms/step)
- Transparent: resize-bench (diagonal): 52 steps in 33213.7ms  (638.72ms/step)

Windows (after):
- Opaque:      resize-bench (diagonal): 52 steps in 8587.4ms   (165.14ms/step)
- Transparent: resize-bench (diagonal): 52 steps in 16579.1ms  (318.83ms/step)

mtmcp_ added on 2026-06-08 17:21:45:

I moved the original caching mechanism up a layer to fix the issue with cache invalidation. This was necessary to prevent the image element from having dependencies on layers above. In this case, the layout calls into the elements to determine if they need to be cached, and caching the pixmap accordingly.

I have attached a new layout-node patch file which is my local diff: fossil diff --verbose --from main --to tip-754-ttk-composed-pixmaps-cache as I don't have commit access here.

I also updated the test script to animate the background color and use a transparent checkerboard pattern on the trough. The results on my side show approximately 200ms shaved off the resize times, and as far as I can tell it should fix the correctness bugs mentioned so far.


mtmcp_ added on 2026-06-04 21:02:34:

Yeah that's probably the best place to start. I'll try to get some profiling results in the coming week.


marc_culler (claiming to be Marc Culler) added on 2026-06-03 17:58:11:
Schelte's report seemed to suggest that all of the "slowness" is from the
blending operation.  That is consistent with your observations since
the caching also eliminates the blending.  So I think we need some
measurements that identify where the time is being spent.  How about
an old-fashioned profiling run?

mtmcp_ added on 2026-06-03 14:21:14:

Marc,

I appreciate the concern about AI usage here. I shouldn't have posted a fragment of the thought process as that sort of derails the conversation without the full context.

However, I would like to speak to the merit of working a solution-space using AI. It is somewhat different and can feel roundabout, like steering a ship and sometimes having to double-back and correct course. It takes some getting used to, but is certainly valid and, contrary to how it might seem, I am reading everything it does and gaining insight into the problem/solution space.

Some things to note, and reasons we shouldn't abandon the caching direction:

  1. The performance issues were exposed due to resizing a window with custom scrollbar elements. However, the current solution should apply to anything that uses image elements.
  2. The AI was correct in many ways, albeit incorrect in others. Such is the nature of a first-draft. The first draft increased speed on Windows & Linux by 40% though (something we noted and everyone helped investigate).
  3. The valuable feedback provided by everyone here helped me to understand the problem space further. I have minimal experience in the Tk internals, so getting stakeholder feedback is crucial to finding the right solution here.
  4. The main outstanding issue as far as I can tell is that the widget changing without a cache invalidation causes stale cache elements.

I went through the motions of implementing the A + B + D and sure enough, the results were no better than without any patch. That was exploring the solution space still and ruling things out, steering in the wrong direction and circling back. That's fine, it was a lot of code I didn't have to write myself courtesy of Claude but I did read through all of it and burn brain cells in the process.

But as my gut instinct was just that C solution, not actually using bounds detection but there's simpler ways to invalidate the cache when dependent widgets change, seems highly likely to fix the stale-cache bug while keeping the 40% performance improvements. It will just expand outward of the ttkImage.c file.

It is also worth noting that the alpha blending is a bottleneck, but it seems to run orthogonal to the algorithmic changes we're discussing. I am thinking both of these solutions could benefit.

Perhaps we will find one-hundred ways not to do it, and then the perfect solution will fall into place. :)

I will explore TkpPutRGBAImage after ruling out the fix for the changing widgets. If we can get 40% performance increase with a correct fix that is a nice place to start. TkpPutRGBAImage will impact much broader surface area, but I am happy to explore that direction.


oehhar added on 2026-06-03 09:03:35:
Marc,
thanks for the post. This is also my knowledge on the MS-Windows side.
In addition, a general solution, and not only for scrollbars, would be great.

Thanks,
Harald

marc_culler (claiming to be Marc Culler) added on 2026-06-02 19:16:45:
One thing about chatbots is that they **love** to chat with you.  An endless
re-design cycle with ever-expanding scope is just dandy with them.  I suggest
that we give the chatbots a rest and think about the core problem with brains,
looking for a pleasing compromise between simplicity and improvement.

Here are a few ideas.

Claude claims "this produces 6,000-12,000 GDI create/destroy cycles per second".
Maybe we could start there.  I have no information about how expensive these
so-called cycles are, but let's assume that is the key issue.  

Looking at TkPutImage in the windows port I see:

    bitmap = (HBITMAP)SelectObject(dcMem, bitmap);
    BitBlt(dc, dest_x, dest_y, (int) width, (int) height, dcMem, src_x, src_y,
            SRCCOPY);
    DeleteObject(SelectObject(dcMem, bitmap));
    DeleteDC(dcMem);
    TkWinReleaseDrawableDC(d, dc, &state);

where the bitmap which appears as an argument to SelectObject was created by
CreateDIBitmap.  I guess Claude is referring to CreateDIBitmap and
DestroyObject as the cycle.

If CreateBitmap and DestroyObject are what is slow, why not cache
the bitmap in TkPutImage, so if TkPutImage were called many times with the
identical image it could avoid all but one of the CreateBitmap and DeleteObject
calls?

A second idea would be: since Schelte finds that the bottleneck is the alpha
blending, rather than the creation and destruction of bitmaps, how about
trying to implement TkpPutRGBAImage for Windows, so that the GPU could be
used to do the alpha blending.  I am sure Windows supports using the GPU
for such a task.

A third idea would be to treat the special case of an image which is 1 pixel
high (in the vertical case), or 1 pixel wide (in the horizontal case), as the
primary application for the tiling construction, and try to optimize for that
case. I believe that using StretchBlt might allow tiling an entire scrollbar
trough with one function call instead of 6000 -12000 calls to BitBlt.

There are probably many other, possibly relatively simple, optimizations which
humans could come up with, if they were to think about it.

mtmcp_ added on 2026-06-02 14:14:49:

Working with Claude Opus 4.8 and the review notes provided another issue surfaced. I will show the summary info here so folks can get a sense of the logic.

The issues

  1. A lower layer in the same widget changed at constant (geometry, state) of the upper element — e.g. an animated lower image element ($photo put …) whose own ImageElementImageChanged invalidates its cache and re-renders it, changing d beneath an upper element whose key is unchanged. The upper element keeps showing the old composite.
  2. Cross-widget contamination. Because ImageData + its single cachedPixmap are shared across all widgets using the style, two scrollbars with identical (width, height, dst.x, dst.y, state) collide on the one cache slot. The key has no widget/window identity. Widget B's translucent element gets blitted with widget A's baked background. (Opaque elements happen to survive this, which is why it hasn't obviously blown up yet — most stock themed elements are opaque.) This is worse than a perf regression; it's wrong pixels, and it's present today, not hypothetical.
  3. Genuinely external content behind a transparent region.Worth being precise here: in Ttk's pipeline each widget renders only its own elements into its own per-frame pixmap d (BeginDrawing → Tk_GetPixmap, fresh each frame, ttkWidget.c:56), and the OS composites separate windows. So a standalone ttk::scrollbar over a separate text widget never sees that text widget's pixels in d — case #3 only materializes inside a single widget's pixmap (compound/overlay widgets that draw live content beneath a translucent element). This is narrower than it sounds, but it's the same root cause as #1.

Solution space

Four directions, roughly in increasing ambition. They compose.

A. Opacity gate (correctness floor, cheapest). Determine once whether the element renders fully opaque over dst (you can derive it from the photo instance's COMPLEX_ALPHA/alpha-table that already exists, or by scanning the source pixels actually used at build time). Only use the cross-frame composite cache when opaque — and when opaque, drop the seed copy. For non-opaque elements, don't keep a cross-frame composite at all. This alone makes the cache correct; it just doesn't accelerate translucent elements.

B. Decouple the cache from the background = "tile once, blend per frame." The GDI thrash comes from ~200 per-tile DIB create/blit/destroy cycles, not from compositing. So cache the tiled element as an RGBA image (one expensive tiling pass, cached by geometry+state), then each frame do a single BlendComplexAlpha + TkPutImage of that whole RGBA buffer onto the live d. You re-read the live background every frame (so it's always correct, cases #1–#3 all handled), but you've still collapsed 200 ops → ~1. This is the correct-and-fast answer for translucent elements and pairs naturally with A (opaque → XCopyArea fast path; translucent → cached-RGBA + one blend).

C. Hierarchy/graph + bounds-intersection invalidation (the approach you gestured at). Within a DrawWidget pass, track the bounding boxes of elements that cache-missed (i.e., changed) this frame. Before letting element E take its fast path, if any already-drawn element's changed-bbox intersects E.dst, force a miss for E. This is pure rectangle intersection — no pixel readback — and, crucially, it is complete for the in-widget case: everything in d came from an element draw this pass, so "did the background under E change?" is exactly "did any lower element re-render under E's bounds?" This cleanly kills case #1. It needs a small amount of per-pass scratch state (a list of dirty rects), which means threading a context through the draw, i.e., the "re-think in terms of the graph" you mentioned.

D. Fix the cache scope (structural, addresses case #2). Move the cache off the shared ImageData. Minimum viable: add window identity to the key (and don't rely on the per-frame d, which is freshly allocated each frame — key on Tk_WindowId, which is stable). Better: a small per-window cache (Ttk elements are intentionally stateless, so this means a side table keyed by (element, window) hung off WidgetCore, or a tiny N-entry LRU on ImageData). Without this, even a correct per-frame algorithm can cross-contaminate two widgets sharing a style.

Claude's recommendation: A + D as the correctness baseline (opacity gate + window-scoped key — small, shippable, makes the current code correct), then B for translucent performance and/or C if you want translucent elements to keep a cross-frame cache. A+D is a modest delta on the existing patch; B and C are the "re-think." Notably, A+B together let you delete the seed-copy hack entirely, which removes the whole baked-background failure mode by construction.


marc_culler (claiming to be Marc Culler) added on 2026-06-02 14:11:30:
Harald,

> What I understood from the patch is that (semi-)transparent pixels
> are always slow, as they are painted pixel by pixel.

That is not correct.

In generic/tkImgPhInstance.c there is conditional code bracketed by
#ifndef TK_CAN_RENDER_RGBA ... #endif which is responsible for blending
an image with transparency over an opaque image.

When TK_CAN_RENDER_RGBA is not defined the function BlendComplexAlpha
is called.  It does the compositing operation on a pixmap in CPU memory.
This blending involves copying a rectangle from the window into a pixmap,
blending the pixmap with the image, using the image's alpha channel, and
then copying the pixmap back to the window.  This all happens in the CPU.

When TK_CAN_RENDER_RGBA is defined, the function TkpPutRGBAImage is used.
That is a platform-specific function which can use the graphics operations
provided by the OS to do the blending.

The only platform which defines TK_CAN_RENDER_RGBA at the moment is macOS.
(Wayland will also use it, as the Wayland port has to interact directly
with the Mesa graphics driver using OpenGL.)  The macOS port draws directly
to the window's backing store CGImage, so it can use Apple's compositing
features to do the blending, and it happens in the GPU.

mtmcp_ added on 2026-06-02 13:38:57:

Thanks everyone for testing this and providing much needed feedback.

Several ideas have surfaced which need more rigorous testing.

  1. The cached image will be stale when a cached, transparent image shows a window underneath any time the window underneath changes. This means the patch needs to be re-thought in terms of the hierarchy/graph and probably needs to be modified to allow a quick invalidation step using bounds detection or something similar.

  2. High-Dpi rendering needs to be investigated further. It seems at the moment the caching works with it, but we should be able to expose this to be certain.

  3. Opaque versus transparent rendering.

  4. Possibly other ways to move the transparency directly into hardware-accelerated rendering. CPU-Side rendering should be okay if the caching works properly but it still seems too slow. In my mind hovering around 100ms refresh would be ideal because it is minimally perceptible by humans.


oehhar added on 2026-06-02 11:39:27:

What I understood from the patch is that (semi-)transparent pixels are always slow, as they are painted pixel by pixel.

The patch reduces the number of semitransparent pixels by cacheing the image, if the image is composed of multiple images which are superposed and have semi-transparent pixels. If there are any pixels not covered by the cached image (and a background shines through) the performance is bad (but less bad than without the patch, if any semi-transparent pixels change status to non-transparent par the overlay procedure).

The fact that semi-transparent pixels are slow (on MS-Windows) is not addressed. Only, the number of those pixels is reduced.

Does this make sense?

Thanks for all, Harald


sbron added on 2026-06-02 10:10:33:

I ran the diagonal sweep test with both Tk 9.0.0 and the tip-754-ttk-composed-pixmaps-cache branch (commit d2bcc173) on linux x11. There doesn't seem to be a significant change in the perf numbers reported. Resizing is horribly slow in both cases, compared to a modified version of the script that doesn't use semi-transparent pixels. In that last case, the branch version is more than twice as fast as 9.0.0 and more than 20 times as fast as the original script.

Original script:

9.0 resize-bench (diagonal): 52 steps in 14624.2ms (281.24ms/step)
tip resize-bench (diagonal): 52 steps in 14684.8ms (282.40ms/step)

Modified script:

8.6 resize-bench (diagonal): 52 steps in 1751.7ms (33.69ms/step)
9.0 resize-bench (diagonal): 52 steps in 1577.5ms (30.34ms/step)
tip resize-bench (diagonal): 52 steps in 694.4ms (13.35ms/step)

This is the diff for the modification I used:

--- test_scrollbar_perf.tcl~    2026-06-02 11:12:38.043915660 +0200
+++ test_scrollbar_perf.tcl     2026-06-02 11:37:10.639595544 +0200
@@ -46,10 +46,10 @@
 img_thumb transparency set 13 0 1
 img_thumb transparency set 0 23 1
 img_thumb transparency set 13 23 1
-img_thumb put {{#80505050}} -to 1 0 2 1
-img_thumb put {{#80505050}} -to 12 0 13 1
-img_thumb put {{#80505050}} -to 1 23 2 24
-img_thumb put {{#80505050}} -to 12 23 13 24
+img_thumb put {{#505050}} -to 1 0 2 1
+img_thumb put {{#505050}} -to 12 0 13 1
+img_thumb put {{#505050}} -to 1 23 2 24
+img_thumb put {{#505050}} -to 12 23 13 24
 
 image create photo img_trough -width 14 -height 14
 img_trough put [string repeat {{#2a2a2a} } 14] -to 0 0 14 14

marc_culler (claiming to be Marc Culler) added on 2026-05-31 21:55:23:
I tested the bugfix branch on Windows with a high-dpi EVO display scaled to
200%.  The test script ran and the scrollbar slider was (barely) visible
as it is on other platforms.  So I take that as empirical proof that there
is not problem with XCopyArea on Windows with a high-dpi display, although
I can't explain what is making it work.

marc_culler (claiming to be Marc Culler) added on 2026-05-31 21:07:46:
I guess there is currently no issue with high-dpi screens in linux, because
the only scaling possible is to change the screen resolution.  For Tk a
pixel in a pixmap is equivalent to a pixel on the screen, and Xorg manages
that via screen resolution.  A default 200x200 Wish window is not
large enough for the title bar to hold the default window controls on a
high-dpi screen, but that is true no matter how you change resolution.
So XCopyArea should work fine in all cases.

marc_culler (claiming to be Marc Culler) added on 2026-05-31 15:17:14:
Hi Mason,

I also worked on the code, after your patch, and added comments to try
to explain what was going on there after doing my best to understand the
situation.  I will commit my changes and you can look at them.

However, I think there are a couple of serious issues with this change
which I do not think have been discussed.

1. The code assumes that the destination rectangle will look the same
whenever the widget containing the image element has the same state. I
don't think that is always a valid assumption.  I think it might be valid
it the widget is opaque, so the background of the destination rectangle
gets completely drawn when drawing the parts of the widget that are drawn
before the destination rectangle is tiled.  But I can imagine cases where
that is not the case.  For example, a scrollbar with a transparent trough,
through which one can see the window content underneath the scrollbar.
That content might change without any change to the scrollbar, for example
if the content was animated.  And in that case the cached pixmap would not
match its surroundings if copied into the widget.

2. It is not clear to me that the patch handles high-dpi screens correctly.
When you create an mxn pixmap it has m*n pixels, each represented by 4
bytes of CPU memory.  But an mxn rectangle on a 2x high-dpi screen has
(2*m)*(2*n) screen pixels, representing m*n logical pixels, stored in GPU
memory. I believe that TkPutImage, because it is  using image drawing code
provided by the OS, can render the mxn pixmap into a (2*m)x(2*n) framebuffer
rectangle by linearly scaling the pixmap image.  I could certainly 
be wrong, but I suspect that XCopyArea does not handle the scaling the
same way.  (I do know that this issue is the key blocker for extending
XCopyArea so that it can work with pixmap source or destination drawables.)

What do you know about #2?  How is the high-dpi rendering handled by the
patched code versus the previous code?

I don't know anything about BitBlt (used by TkPutImage on Windows) but
Gemini told me that:
 "The operation is ultimately translated into a GPU instruction or a
  high-speed contiguous chunk transfer (such as a hardware-accelerated
  memcpy) via the Direct3D/DXGI stack or the Desktop Window Manager".
That sort of sounds like a Direct3D version of glBlitBuffer, which
automatically linearly scales an image if the source and destination
rectangles are different sizes.  But it apparently can have a CPU
source.  On the other hand Gemini says that BitBlt does not do
scaling and that one would have to use StretchBlt for that.  In fact,
StretchBlt is used in GdiImage, but it looks like that is only used by
the Canvas widget, not by XPutImage.

And then there is the question of how the XCopyArea and TkPutImage compare
under X11 with a high-dpi screen.

So I really don't know where things stand regarding high-dpi screens.

While I may be making mountains out of molehills, I would feel much better
about this if some of these concerns were addressed by someone who knows
more about the situation than I do.

oehhar added on 2026-05-31 15:06:17:

Mason, I gave you commit rights on the tk repository. Could you update the branch?

Thanks for all, Harald


mtmcp_ added on 2026-05-31 14:08:48:

Marc,

I went ahead and made the changes to use TK_NO_DOUBLEBUFFERING instead and replaced the comment "blit" with "copy". Using the define as you requested is cleaner and I am happy with that.

The latest changes are in the patch file: ttk-image-macos-fixes.diff.


mtmcp_ added on 2026-05-31 13:46:06:

Marc,

I agree re. the blitting comments. I will have them removed.

For the TK_NO_DOUBLEBUFFERING define, it was my intention to make more explicit what is happening. I prefer the most self-documenting code if possible. It is harder to interpret double-negatives like #ifndef TK_NO_DOUBLEBUFFERING as opposed to #ifdef TK_CAN_XCOPYAREA_PIXMAP. Also, it seemed to me that we are not removing the capability due to lack of double-buffering, we are in fact removing it because macOS (at the moment) does not support XCopyArea for Pixmaps. So, this seems more correct. There is no harm in adding a define in this one C file as it won't impact anything else. That being said, if you insist I am happy to switch it around. :)

I ran a test using the existing script and inferred that it is as it was without the patch because the performance is virtually identical on macOS before/after the patch on macOS. As you stated prior the "improved performance" before was comically because it wasn't doing anything useful at all!

However, I am going to do a bit more digging just to make sure we get better coverage. Actually, I am thinking this could be made a unit test if we'd like. I can revise the script for automation it could be a nice regression check.


marc_culler (claiming to be Marc Culler) added on 2026-05-30 23:35:43:
Never mind -- I see that I can use the performance script to verify
that the scrollbar looks like a scrollbar, and assume that it looks
the same as you intended.

marc_culler (claiming to be Marc Culler) added on 2026-05-30 21:09:07:
Mason, Do you have a script that I could run to make sure that your
custom scrollbar appears correctly on macOS?

marc_culler (claiming to be Marc Culler) added on 2026-05-30 20:59:33:
I don't think we need to add a new #define constant which we plan to remove
as soon as possible.  Why not use TK_NO_DOUBLEBUFFERING which is what avoids
essentially the same use of XCopyArea throughout the generic code?  It also
probably represents what is going on here; that is to say that the fact
that macOS draws directly to the screen was probably a big part of why
it did not have the big delays in the first place.

Also, Claude seems very determined to characterize XCopyArea as a "blit" but
I think that is wrong.  You cannot blit from a Pixmap (which lives in CPU
memory) to the screen.  That is always a slow copy operation.  A blit moves
pixel data efficiently between two framebuffers, both living in GPU memory.
I think it is misleading and unhelpful to refer to XCopyArea as a "blit".
Could we please remove that comment?

mtmcp_ added on 2026-05-30 16:57:52:

I have attached the fix so you may test until I can push it (or if you'd like to commit it): ttk-image-macos-fixes.diff

It includes a couple other correctness fixes like zeroing variables and using Tcl_Alloc/Free to be consistent.


mtmcp_ added on 2026-05-30 16:46:53:
Marc,

Thank you for finding that issue. I have completed a fix but I don't seem to have write access to the Tk branch.

marc_culler (claiming to be Marc Culler) added on 2026-05-29 21:22:30:
>> I retry a CI run. If this succeeds, I will merge to main.
>> No merge to core-9-0-branch. 

Harald,

Please do not merge to main until the cache code is made conditional on #ifndef NO_TK_DOUBLE_BUFFERING
It will break macOS. Please don't do that.

marc_culler (claiming to be Marc Culler) added on 2026-05-29 16:08:11:
IMPORTANT!

I juat remembered why I was so worried about this patch for macOS.  It
woke me up in the middle of the night.  It has a fatal flaw, for macOS
only: It will not render the cached Pixmaps!  The fact that the test
script runs faster is meaningless if the drawing is not happening.

The reason that the drawing will hot happen has two parts:
(1) The drawing code includes the following line:

XCopyArea(Tk_Display(tkwin), imageData->cachedPixmap, d, gc,
		0, 0, (unsigned)dst.width, (unsigned)dst.height, dst.x, dst.y);

(2) The macOS port of XCopyArea is currently incomplete and does not
support a source or destination drawable which is a Pixmap.  If passed
a Pixmap for either it just returns BadDrawable and does no drawing.
(See lines 1125-1128 in macosx/TkMacOSXImage.c.)

While XCopyArea is used in the generic code to do makeshift
double-buffering (copy from window to pixmap, draw on pixmap, copy
back to window), that code is conditional, based on the compiler
constant TK_NO_DOUBLE_BUFFERING, which is defined by the macOS port only.

Finishing XCopyArea for macOS is an important project which needs to be
done.  But it is non-trivial for a few reasons, a major one being that
it needs to deal with Retina displays.  I have been thinking about this
lately, after it came up in the ticket [4af5ca1921].  I think the simplest
approach would be to require that pixmaps are always using logical pixels
and let the OS do the linear scaling to the 2x2 pixels used by a Retina
display.  That will mean that pixmaps displayed on a Retina display are
more pixelated than the origianl and will not be anti-aliased as well. I
don't see a practical alternative to that at the moment, but I am not
sure it will be adequate for projects like the one discussed in
[4af5ca1921].

Another aspect of this is that the scaling which would be required if
XCopyArea always treats pixmaps as consisting of logical pixels is that
that scaling makes rendering considerably slower.  So if the test script
were actually rendering the pixmaps, instead of just checking that they
are pixmaps and returning BadDrawable when the answer is yes, it might
well mean that the macOS speed up would be considerably less, at least
on Retina displays.

In any case, I think the only reasonable way for this ticket to move
forward without breaking the graphics on macOS would be to make the new
caching code conditional on ifndef(TK_NO_DOUBLE_BUFFERING), which would
exclude macOS, and to revert to the original code when
TK_NO_DOUBLE_BUFFERING is defined.

I am also curious about whether the original Windows problem is affected
by whether the computer has a high-DPI screen.

oehhar added on 2026-05-29 09:18:34:

Ok, thanks. Now, I understood the tests. You have to resize the window, not scroll the scrollbar. Improvement is directly visible on Windows, thanks for that.

I retry a CI run. If this succeeds, I will merge to main. No merge to core-9-0-branch.

Any opinions on those proposals?

Thanks for all, Harald


marc_culler (claiming to be Marc Culler) added on 2026-05-28 18:23:06:
On an M3 macbook Air (with Retina display, which affects XCopyArea) running
macOS 26.4 I got:

Pre-patch: resize-bench (diagonal): 52 steps in 2980.8ms  (57.32ms/step)
Post-patch: resize-bench (diagonal): 52 steps in 1970.2ms  (37.89ms/step)

There were no crashes!

The non-regression test failures were the usual ones that I see on that
system (event-9.18 and event-0.19) plus two new ones (send-13.1 and send-13.2)
which I am sure are unrelated.

So I would say that the patch looks like a significant improvement for
macOS.

mtmcp_ added on 2026-05-28 17:38:19:
To summarize the performance changes below, I am getting about 40% improvement on Windows/Linux and 20% improvement for macOS.

Ideally, the Windows/Linux systems should have sub-100 ms latency on the resizing. Perhaps a follow-up change can be done to improve these systems. However, this minimally-invasive change seems to provide improvements to performance across the board.

mtmcp_ added on 2026-05-28 17:17:32:

Some macOS performance stats also show an improvement in performance after applying the patch:

macOS (arm64):
- before: "resize-bench (diagonal): 52 steps in 2785.8ms  (53.57ms/step)"
- after:  "resize-bench (diagonal): 52 steps in 2232.5ms  (42.93ms/step)"

mtmcp_ added on 2026-05-28 16:15:16:

I have updated the performance test script so that it captures performance to quantify the results of the change. Here is my local results on Windows and Linux (WSL2):

Linux (WSL2):
- before: "resize-bench (diagonal): 52 steps in 17462.2ms  (335.81ms/step)"
- after:  "resize-bench (diagonal): 52 steps in  9890.8ms  (190.21ms/step)"

Windows:
- before: "resize-bench (diagonal): 52 steps in 18896.7ms  (363.40ms/step)"
- after:  "resize-bench (diagonal): 52 steps in 11025.1ms  (212.02ms/step)"

oehhar added on 2026-05-28 09:46:43:

The comments so far:

Schelte:

I could not detect any difference in my testing. But that also means I found no problems with the change.

Mason:

When running the performance test script linked at the bottom of TIP 754, in order to see the lag without the patch, one must repeatedly drag the diagonal grip to resize the window. On Windows, the perceptible lag is quite noticeable without the patch. However, ideally this should be quantified so I can update the test script to provide performance metrics.

Marc:

What bothers me is that the problem seems to only be for Windows but the solution affects other platforms and I don't think that it has been carefully tested on, say, macOS, to see whether it improves or degrades performance. The macOS platform does not use any pixmaps when drawing to a window, and I think that may be an important difference.


Attachments: