|
2026-09-29
| ||
| 19:32 | • Closed ticket [32f2e19d]: Windows: photo -gamma and -palette are ignored since RGBA rendering plus 9 other changes artifact: 53048d6f user: oehhar | |
|
2026-06-17
| ||
| 13:42 | [7caf9e9e] Cache Composed Pixmaps for Ttk Image Elements check-in: 0f84a33e user: mtmcp_ tags: trunk, main | |
| 06:30 | • Ticket [7caf9e9e] Cache Composed Pixmaps for Ttk Image Elements status still Closed with 6 other changes artifact: 0e104f42 user: oehhar | |
|
2026-06-16
| ||
| 20:26 | • Ticket [7caf9e9e]: 5 changes artifact: f4299dbd user: mtmcp_ | |
| 08:59 | • Ticket [7caf9e9e]: 5 changes artifact: 779cdc29 user: oehhar | |
| 08:58 | • Ticket [7caf9e9e]: 6 changes artifact: 7e86b638 user: oehhar | |
| 01:13 | • Ticket [7caf9e9e]: 4 changes artifact: e89b8472 user: mtmcp_ | |
|
2026-06-15
| ||
| 17:28 | • Ticket [7caf9e9e]: 5 changes artifact: 7f8e2db2 user: mtmcp_ | |
| 08:38 | • Ticket [7caf9e9e]: 5 changes artifact: 20ede9f8 user: oehhar | |
| 08:24 | • Closed ticket [7caf9e9e]. artifact: 6146168f user: oehhar | |
| 08:21 | [7caf9e9e] Cache Composed Pixmaps for Ttk Image Elements check-in: 3bbe54bc user: oehhar tags: trunk, main | |
|
2026-06-14
| ||
| 12:19 | • Ticket [7caf9e9e] Cache Composed Pixmaps for Ttk Image Elements status still Open with 4 other changes artifact: 0ec3bfac user: oehhar | |
|
2026-06-13
| ||
| 19:34 | • Ticket [7caf9e9e]: 3 changes artifact: 3da3d8d4 user: sbron | |
| 18:11 | • Ticket [7caf9e9e]: 3 changes artifact: 72ca9c70 user: mtmcp_ | |
| 11:35 | • Add attachment test_scrollbar_perf.tcl to ticket [7caf9e9e] artifact: 09fdf631 user: mtmcp_ | |
|
2026-06-12
| ||
| 08:09 | • Ticket [7caf9e9e] Cache Composed Pixmaps for Ttk Image Elements status still Open with 4 other changes artifact: 0a1c767a user: oehhar | |
|
2026-06-11
| ||
| 16:29 | • Ticket [7caf9e9e]: 3 changes artifact: 664192a3 user: erikleunissen | |
| 13:28 | • Ticket [7caf9e9e]: 3 changes artifact: 6d34fc1c user: mtmcp_ | |
| 12:29 | • Add attachment test_sections.patch to ticket [7caf9e9e] artifact: 23e9c1a0 user: erikleunissen | |
| 12:27 | • Ticket [7caf9e9e] Cache Composed Pixmaps for Ttk Image Elements status still Open with 3 other changes artifact: 105578fe user: erikleunissen | |
| 11:07 | • Ticket [7caf9e9e]: 3 changes artifact: 884d94aa user: mtmcp_ | |
| 08:53 | • Ticket [7caf9e9e]: 4 changes artifact: 2d956c61 user: oehhar | |
| 03:25 | • Ticket [7caf9e9e]: 3 changes artifact: 1a516ad1 user: mtmcp_ | |
|
2026-06-10
| ||
| 07:36 | • Ticket [7caf9e9e]: 4 changes artifact: a117b87a user: oehhar | |
|
2026-06-09
| ||
| 16:10 | • Ticket [7caf9e9e]: 3 changes artifact: e0582294 user: mtmcp_ | |
| 16:07 | • Add attachment tk-9.0.3-ttk-image-el-cache.diff to ticket [7caf9e9e] artifact: c9e16e1c user: mtmcp_ | |
| 15:00 | • Ticket [7caf9e9e] Cache Composed Pixmaps for Ttk Image Elements status still Open with 3 other changes artifact: a7b8140b user: mtmcp_ | |
| 14:38 | • Ticket [7caf9e9e]: 3 changes artifact: 086058e0 user: mtmcp_ | |
| 14:19 | • Add attachment ttk-layout-node-cache.diff to ticket [7caf9e9e] artifact: d3b6748d user: mtmcp_ | |
| 13:20 | • Ticket [7caf9e9e] Cache Composed Pixmaps for Ttk Image Elements status still Open with 4 other changes artifact: a60cb599 user: oehhar | |
| 01:55 | • Ticket [7caf9e9e]: 3 changes artifact: e0893df7 user: mtmcp_ | |
| 01:45 | • Add attachment test_scrollbar_perf.tcl to ticket [7caf9e9e] artifact: c8dd7346 user: mtmcp_ | |
|
2026-06-08
| ||
| 17:30 | • Delete attachment "ttk-image-macos-fixes.diff" from ticket [7caf9e9e] artifact: dd8c66df user: mtmcp_ | |
| 17:21 | • Ticket [7caf9e9e] Cache Composed Pixmaps for Ttk Image Elements status still Open with 3 other changes artifact: 53825d16 user: mtmcp_ | |
| 17:09 | • Add attachment ttk-layout-node-cache.diff to ticket [7caf9e9e] artifact: d5bb5659 user: mtmcp_ | |
| 16:45 | • Add attachment test_scrollbar_perf.tcl to ticket [7caf9e9e] artifact: 045fc6fa user: mtmcp_ | |
|
2026-06-04
| ||
| 21:02 | • Ticket [7caf9e9e] Cache Composed Pixmaps for Ttk Image Elements status still Open with 3 other changes artifact: 573f5bf4 user: mtmcp_ | |
|
2026-06-03
| ||
| 17:58 | • Ticket [7caf9e9e]: 4 changes artifact: 8f10d2e9 user: marc_culler | |
| 14:21 | • Ticket [7caf9e9e]: 3 changes artifact: 63d24c2a user: mtmcp_ | |
| 09:03 | • Ticket [7caf9e9e]: 4 changes artifact: 413ac4e2 user: oehhar | |
|
2026-06-02
| ||
| 19:16 | • Ticket [7caf9e9e]: 4 changes artifact: b31d63eb user: marc_culler | |
| 14:14 | • Ticket [7caf9e9e]: 3 changes artifact: 439f74d2 user: mtmcp_ | |
| 14:11 | • Ticket [7caf9e9e]: 4 changes artifact: 7d94c3dc user: marc_culler | |
| 13:38 | • Ticket [7caf9e9e]: 3 changes artifact: 920084ae user: mtmcp_ | |
| 11:39 | • Ticket [7caf9e9e]: 4 changes artifact: c8ddbe33 user: oehhar | |
| 10:10 | • Ticket [7caf9e9e]: 3 changes artifact: ba9ea865 user: sbron | |
|
2026-05-31
| ||
| 21:55 | • Ticket [7caf9e9e]: 4 changes artifact: 026981f0 user: marc_culler | |
| 21:07 | • Ticket [7caf9e9e]: 4 changes artifact: cafd8966 user: marc_culler | |
| 15:17 | • Ticket [7caf9e9e]: 4 changes artifact: 8a40fdf5 user: marc_culler | |
| 15:06 | • Ticket [7caf9e9e]: 4 changes artifact: aa174625 user: oehhar | |
| 14:08 | • Ticket [7caf9e9e]: 3 changes artifact: d19270f4 user: mtmcp_ | |
| 14:05 | • Add attachment ttk-image-macos-fixes.diff to ticket [7caf9e9e] artifact: 5758fb0f user: mtmcp_ | |
| 13:46 | • Ticket [7caf9e9e] Cache Composed Pixmaps for Ttk Image Elements status still Open with 3 other changes artifact: 451a72f2 user: mtmcp_ | |
|
2026-05-30
| ||
| 23:35 | • Ticket [7caf9e9e]: 4 changes artifact: 0581cda8 user: marc_culler | |
| 21:09 | • Ticket [7caf9e9e]: 4 changes artifact: 2af3bea4 user: marc_culler | |
| 20:59 | • Ticket [7caf9e9e]: 4 changes artifact: 176e1a9f user: marc_culler | |
| 16:57 | • Ticket [7caf9e9e]: 3 changes artifact: ee1cecc2 user: mtmcp_ | |
| 16:54 | • Add attachment ttk-image-macos-fixes.diff to ticket [7caf9e9e] artifact: 3012d15c user: mtmcp_ | |
| 16:46 | • Ticket [7caf9e9e] Cache Composed Pixmaps for Ttk Image Elements status still Open with 3 other changes artifact: 33c32042 user: mtmcp_ | |
|
2026-05-29
| ||
| 21:22 | • Ticket [7caf9e9e]: 4 changes artifact: 6ef42d2f user: marc_culler | |
| 16:08 | • Ticket [7caf9e9e]: 4 changes artifact: 893226e6 user: marc_culler | |
| 09:18 | • Ticket [7caf9e9e]: 4 changes artifact: 057e9a03 user: oehhar | |
|
2026-05-28
| ||
| 18:23 | • Ticket [7caf9e9e]: 4 changes artifact: 089302fb user: marc_culler | |
| 17:38 | • Ticket [7caf9e9e]: 3 changes artifact: b7c0337c user: mtmcp_ | |
| 17:17 | • Ticket [7caf9e9e]: 3 changes artifact: 3a12067f user: mtmcp_ | |
| 16:18 | • Ticket [7caf9e9e]: 3 changes artifact: 933d9a86 user: mtmcp_ | |
| 16:15 | • Ticket [7caf9e9e]: 3 changes artifact: 1f94c2a7 user: mtmcp_ | |
| 16:07 | • Add attachment test_scrollbar_perf.tcl to ticket [7caf9e9e] artifact: 617affc5 user: mtmcp_ | |
| 09:46 | • Ticket [7caf9e9e] Cache Composed Pixmaps for Ttk Image Elements status still Open with 4 other changes artifact: 551331e1 user: oehhar | |
| 09:33 | • Add attachment test_scrollbar_perf.tcl to ticket [7caf9e9e] artifact: d963a3c1 user: oehhar | |
| 09:28 | • New ticket [7caf9e9e] Cache Composed Pixmaps for Ttk Image Elements. artifact: 2ce66340 user: oehhar | |
| Ticket UUID: | 7caf9e9edcfbee425264b76be61c973c7bd2db1 | |||||||||||||
| Title: | Cache Composed Pixmaps for Ttk Image Elements | |||||||||||||
| Type: | RFE | Version: | main | |||||||||||
| Submitter: | oehhar | Created on: | 2026-05-28 09:28:36 | |||||||||||
| Subsystem: | 41. Photo Images | Assigned To: | oehhar | |||||||||||
| Priority: | 5 Medium | Severity: | Minor | |||||||||||
| Status: | Closed | Last Modified: | 2026-06-17 06:30:13 | |||||||||||
| Resolution: | Fixed | Closed By: | oehhar | |||||||||||
| Closed on: | 2026-06-17 06:30:13 | |||||||||||||
| Description: |
Custom Ttk scrollbar elements created with All changes are confined to RationaleThe On Windows, each During scrolling or window resizing with a custom image-based theme, this produces 6,000-12,000 GDI create/destroy cycles per second, causing visible UI lag and high CPU usage. The same elements render without issue on macOS, which uses hardware-accelerated Core Graphics compositing and avoids per-tile object allocation. Built-in themes like SpecificationPer-element pixmap cacheEach Position (x, y) is included in the key because the cache pixmap is seeded with the destination background content before tiling. This ensures correct alpha compositing for images with semi-transparent pixels — without the background seed, transparent areas would composite against an uninitialized (black) pixmap. If the element's position changes, the underlying background may differ, so the cache must be invalidated. HistoryThis was first proposed as TIP 751 and transfered to a ticket. On each
The cache is invalidated (pixmap freed, dimensions zeroed) when:
If the window is not yet realized ( Off-by-one fix in Ttk_FillThe tiling loop at line 231 uses Properties
Reference ImplementationImplementation is in branch tip-754-ttk-composed-pixmaps-cache. See the attached performance test script: test_scrollbar_perf.tcl for reproducing the issue. Tested with Tcl/Tk 9.0.3 on Windows using custom image-based scrollbar themes. Performance improvement is on the order of 10-50x reduction in GDI object operations during scrolling and resize. | |||||||||||||
| User Comments: |
oehhar added on 2026-06-17 06:30:13:
Great ! If CI is clean, you may merge to main. Thanks for all, Harald mtmcp_ added on 2026-06-16 20:26:16: I re-opened the branch and updated with the following changes:
The third change I think was necessary because this procedure was very confusing to read, so extracting the parts into helper functions improves it. Nothing else sticks out to me, so I am okay with the code for now perhaps until I walk away and come back with fresh eyes during beta tests. Thanks everyone for the prompt feedback and help getting to the root cause on these performance changes. I am excited for the 9.1 beta. :) oehhar added on 2026-06-16 08:59:02: You can also reuse the ticket. Just put comments here. Even if it is closed, you may still edit it. oehhar added on 2026-06-16 08:58:02: Great, thank you. You may also reopen the branch and continue to use it. It is up to you. Click on the commit number, press "edit" and select the cancel "closed" tag option. Thanks for all, Harald mtmcp_ added on 2026-06-16 01:13:36: I see the branch is closed so that answers my question! I will create a new branch and ticket for the updates. :) mtmcp_ added on 2026-06-15 17:28:10: Harald, Yes good point on the documentation. Previously that xrender library was pulled in as a transient dependency from Xft, but now that xrender is used for the Linux TkpPutRGBAImage. I would like to do a couple follow-up changes, for example, there are a few "ckalloc" &co. function calls that should probably be their Tcl_Alloc analogues. I can work the documentation change in. Should I make these changes in the existing branch, or create a new one? Also should I create a new ticket? Thanks, Mason oehhar added on 2026-06-15 08:38:04: The patch introduces the following new dependencies:
There is a new configure switch for Linux: --enable-xrender (default: yes) Is there additional documentation needed ? Thanks for all, Harald oehhar added on 2026-06-15 08:24:08: CI is clean. Schelte is in favor. Merged to main by commit [3bbe54bc]. @Mason: thanks for the awesome work of this very long standing annoyance! Bug closed. Take care, Harald oehhar added on 2026-06-14 12:19:42: Great. Activated CI on head node. When this is clean, we will erge to main. Mason, you may also adopt the main text of this ticket, if the original text changed. Thanks for all, Harald sbron added on 2026-06-13 19:34:54: Amazing. The transparent timings have improved to be on par with the opaque timings on Tk 9.0. And the new opaque resize is almost too fast for the eye to track. Excellent job! Transparent: resize-bench (diagonal): 52 steps in 2378.4ms (45.74ms/step) Opaque: resize-bench (diagonal): 52 steps in 295.0ms (5.67ms/step) mtmcp_ added on 2026-06-13 18:11:27: Okay, Linux TkpPutRGBAImage support was added along with batching (all platforms) to further increase performance. Here are the numbers across the board, all safely under 100ms threshold:
The caching unit tests now gated with "notAqua" to prevent running them on macOS. oehhar added on 2026-06-12 08:09:04: Info from the core list: SchelteIn the ticket, Mason reported that transparent and opaque runs have virtually identical runtime on Windows. On linux I still see a huge difference between the two runs, with transparent being more than 10 times slower than opaque. But still the transparent run is similar to the time Mason reported and I'm running this on quite an old machine. So maybe the opaque run is just really slow on Windows compared to linux? These are my results: Transparent: resize-bench (diagonal): 52 steps in 8590.6ms (165.20ms/step) Opaque: resize-bench (diagonal): 52 steps in 672.6ms (12.93ms/step) In any case, both are significant improvements compared to running the same test with Tk 9.0: Transparent: resize-bench (diagonal): 52 steps in 15718.8ms (302.28ms/step) Opaque: resize-bench (diagonal): 52 steps in 2402.1ms (46.19ms/step) MasonThanks for testing the change out and provide feedback. The gap on Linux is expected at the moment due to the missing TkpPutRGBAImage implementation. So, the gains on Linux are primarily from the caching behavior. On Windows the performance issue was a huge bottleneck to custom theming. Since we can now leverage the nice SVG capabilities in Tk9, it seemed most important to bring Windows down around 100ms (from 500/600ish) in the diagonal resize test with alpha-composited SVG elements. If folks here review the changes and like how it's all going, I will follow up with the Linux TkpPutRGBAImage. I didn't prioritize this one yet because it was already floating in the 100's on the alpha-compositing anyway, but surely we can boost it a lot more. :) Harald
thanks for attacking Linux to. I personally would love, that we first could merge the current branch and then do the next step.
But Wizards are always fast, that is life...
The CI run of today is here:
https://github.com/tcltk/tk/actions?query=
Workflow core-tip-754-ttk-composed-pixmaps-cache
It is the older commit without the test reform by Eric.
We see a failure on MacOS:
The nodeCache tests failed.
https://github.com/tcltk/tk/actions/runs/27400538119/job/80977346594
==== nodecache-1.1 Redraw with nothing changed is served from the cache FAILED
==== Contents of test case:
set ncLog {}
ncExpose .nc
llength $ncLog
---- Result was:
6
---- Result should have been (exact matching):
0
==== nodecache-1.1 FAILED
==== nodecache-1.2 Repeated redraws stay cached FAILED
==== Contents of test case:
set ncLog {}
ncExpose .nc
ncExpose .nc
ncExpose .nc
llength $ncLog
---- Result was:
18
---- Result should have been (exact matching):
0
==== nodecache-1.2 FAILED
==== nodecache-2.3 Redraw after a state change hits again FAILED
==== Contents of test case:
set ncLog {}
ncExpose .nc
llength $ncLog
---- Result was:
6
---- Result should have been (exact matching):
0
==== nodecache-2.3 FAILED
==== nodecache-3.3 Redraw after a resize hits again FAILED
==== Contents of test case:
set ncLog {}
ncExpose .nc
llength $ncLog
---- Result was:
15
---- Result should have been (exact matching):
0
==== nodecache-3.3 FAILED
==== nodecache-5.2 Content change above leaves the node below cached FAILED
==== Contents of test case:
set ncLog {}
nc.over changed 0 0 30 15 30 15
ncExpose .nc
list [ncLogged nc.under] [ncLogged nc.over]
---- Result was:
1 1
---- Result should have been (exact matching):
0 1
==== nodecache-5.2 FAILED
==== nodecache-8.1 Stable parcels stay cached through a layout sweep FAILED
==== Contents of test case:
set ncLog {}
set x0 [winfo x .sw.nc]
for {set w 120} {$w <= 220} {incr w 20} {
.sw.holder configure -width $w
update
}
set x1 [winfo x .sw.nc]
list [expr {$x1 > $x0}] [llength $ncLog]
---- Result was:
1 36
---- Result should have been (exact matching):
1 0
==== nodecache-8.1 FAILED
Perhaps, the tests are not valid on MacOS as the cache is not used. Then, they should be flagged as "not on MacOS".
Could you look into this ?
erikleunissen added on 2026-06-11 16:29:00: Thanks for explaining the choice to create a new test file imgPhInstance.test. I'm entirely convinced by your arguments. Wise decision. As for the image cleanup: I believe there is a "testutils forget image" missing in the section TESTFILE CLEANUP of both test files. mtmcp_ added on 2026-06-11 13:28:12: Erik, Thank you for reviewing the changes. Per your suggestions I have applied the patch and also extended them to nodecache.test. Regarding the new imgPhInstance.test, the file is named for tkImgPhInstance.c per the README convention — these tests exercise the display path (TkImgPhotoDisplay's three tiers), not the photo model API that imgPhoto.test covers and maps function-by-function in its header; tkImgPhInstance.c previously had no coverage, and the display tests carry unusual environment constraints (live screen readback, unobscured TrueColor window) that are better isolated in a file that can be excluded as a unit. erikleunissen added on 2026-06-11 12:27:24: I've had a look at the new test files imgPhInstance.test and ttk/nodecache.test, mostly from the point of view of maintainability. Overall: they look fine to me in this respect ! A few recommendations/suggestions/questions: * Just being curious: why create a new test file imgPhInstance.test instead of adding the new tests to the existing test file imgPhoto.test? * To conform to the usage of comments/headers for standard test sections in other test files, I'd suggest a few re-arrangements. Please see the attached patch. * For the cleanup of images, the Tk test suite already provides some standard procs for test authors in testutils.tcl: imageInit, imageFinish, imageCleanup and friends. You may find these useful instead of manually enumerating image names for cleanup. See also their usage in imgPhoto.test. Regards, Erik Leunissen. mtmcp_ added on 2026-06-11 11:07:06: Harald, The changes in Marc's branch are based upon the old code and we probably should keep the branch closed as many things have diverged. For example, the caching mechanism is moved to the layout and is no longer in the image element directly. In other words, there is no resolution in the merge it will merely be to abandon the older changes (same as closing the branch I think). oehhar added on 2026-06-11 08:53:04: The branch [tip-754-ttk-composed-pixmaps-cache] had two open leaves:
I have closed the branch [d2bcc17325] by Marc to avoid the conflict. Better would be to merge in and to resolve the issues. I tried but was not able to do so. So, here are the commands:
fossil merge d2bcc17325 MERGE generic/ttk/ttkImage.c ***** 4 merge conflicts in generic/ttk/ttkImage.c WARNING: 1 merge conflicts Resolve the merge conflicts and commit with "closed fork". Thanks and sorry, Harald mtmcp_ added on 2026-06-11 03:25:09: Okay, I added a few unit tests to wrap this up and it's ready for review. The unit tests exposed an issue with the cache always being marked dirty for transparent runs due to the background/border invalidating everything on top. After fixing this issue, transparent AND opaque runs have virtually identical runtime, and it is close enough to the 100ms so I am really happy with the results:
I think it is ready to go now that it has the test coverage, pending some final reviews, tests and whatnot. To run the unit tests using nmake:
oehhar added on 2026-06-10 07:36:27: Thanks, great ! It would be great to focus on the main branch so this change goes into the first beta of tk9.1 expected end of the month. You can ask for testing to make it work on all platforms. In addition, CI testing may be triggered if interested. Thanks, Harald mtmcp_ added on 2026-06-09 16:10:46: I added a regression patch for 9.0.3 here (tk-9.0.3-ttk-image-el-cache.diff). Strangely, this runs much faster than my previous nmake build. These results are really good: Windows/MSYS2 (Tk 9.0.3 with TkpPutRGBAImage + cache) - Opaque: resize-bench (diagonal): 52 steps in 4657.0ms ( 89.56ms/step) - Transparent: resize-bench (diagonal): 52 steps in 6920.1ms (133.08ms/step) mtmcp_ added on 2026-06-09 15:00:14: Harald, Sorry for the confusion. It looks like I just had to update my fossil remote. I pushed the latest changes to mtmcp_ added on 2026-06-09 14:38:00: Okay, so I implemented TkpPutRGBAImage on Windows (see latest ttk-layout-node-cache.diff). With that change plus the layout-node caching mechanism, the results seem pretty solid! The gap between opaque and transparent is ~ 40ms, and the overall cost is dramatically less than without this patch. Horizontal and vertical resize are better than diagonal, understandably, so the overall performance improvement with this change will be really nice on Windows, and Linux given regular usage (not the benchmark script). Once macOS can take advantage of the caching it will likely see a speedup as well. Windows (with TkpPutRGBAImage + cache): Diagonal: - Opaque: resize-bench (diagonal): 52 steps in 6827.6ms (131.30ms/step) - Transparent: resize-bench (diagonal): 52 steps in 9053.8ms (174.11ms/step) Horizontal: - Opaque: resize-bench (horizontal): 52 steps in 5495.4ms (105.68ms/step) - Transparent: resize-bench (horizontal): 52 steps in 7858.9ms (151.13ms/step) Vertical: - Opaque: resize-bench (vertical): 42 steps in 3120.4ms (74.30ms/step) - Transparent: resize-bench (vertical): 42 steps in 4941.8ms (117.66ms/step) oehhar added on 2026-06-09 13:20:43: You should have commit rights since some weeks... If not, E-Mail me... Harald mtmcp_ added on 2026-06-09 01:55:14: The test script now allows testing fully opaque and also transparency. Here are current performance metrics with the latest caching patch. As is evident here, the caching dramatically improves the speed for both fully opaque and transparent image elements. This is without any optimization on the blending path, which should of course be subject to the profiling runs. Linux (before): - Opaque: resize-bench (diagonal): 52 steps in 3410.8ms ( 65.59ms/step) - Transparent: resize-bench (diagonal): 52 steps in 20413.2ms (392.56ms/step) Linux (after): - Opaque: resize-bench (diagonal): 52 steps in 1773.3ms ( 34.10ms/step) - Transparent: resize-bench (diagonal): 52 steps in 11097.4ms (213.41ms/step) Windows (before): - Opaque: resize-bench (diagonal): 52 steps in 25325.8ms (487.04ms/step) - Transparent: resize-bench (diagonal): 52 steps in 33213.7ms (638.72ms/step) Windows (after): - Opaque: resize-bench (diagonal): 52 steps in 8587.4ms (165.14ms/step) - Transparent: resize-bench (diagonal): 52 steps in 16579.1ms (318.83ms/step) mtmcp_ added on 2026-06-08 17:21:45: I moved the original caching mechanism up a layer to fix the issue with cache invalidation. This was necessary to prevent the image element from having dependencies on layers above. In this case, the layout calls into the elements to determine if they need to be cached, and caching the pixmap accordingly. I have attached a new layout-node patch file which is my local diff: I also updated the test script to animate the background color and use a transparent checkerboard pattern on the trough. The results on my side show approximately 200ms shaved off the resize times, and as far as I can tell it should fix the correctness bugs mentioned so far. mtmcp_ added on 2026-06-04 21:02:34: Yeah that's probably the best place to start. I'll try to get some profiling results in the coming week. marc_culler (claiming to be Marc Culler) added on 2026-06-03 17:58:11: Schelte's report seemed to suggest that all of the "slowness" is from the blending operation. That is consistent with your observations since the caching also eliminates the blending. So I think we need some measurements that identify where the time is being spent. How about an old-fashioned profiling run? mtmcp_ added on 2026-06-03 14:21:14: Marc, I appreciate the concern about AI usage here. I shouldn't have posted a fragment of the thought process as that sort of derails the conversation without the full context. However, I would like to speak to the merit of working a solution-space using AI. It is somewhat different and can feel roundabout, like steering a ship and sometimes having to double-back and correct course. It takes some getting used to, but is certainly valid and, contrary to how it might seem, I am reading everything it does and gaining insight into the problem/solution space. Some things to note, and reasons we shouldn't abandon the caching direction:
I went through the motions of implementing the A + B + D and sure enough, the results were no better than without any patch. That was exploring the solution space still and ruling things out, steering in the wrong direction and circling back. That's fine, it was a lot of code I didn't have to write myself courtesy of Claude but I did read through all of it and burn brain cells in the process. But as my gut instinct was just that C solution, not actually using bounds detection but there's simpler ways to invalidate the cache when dependent widgets change, seems highly likely to fix the stale-cache bug while keeping the 40% performance improvements. It will just expand outward of the It is also worth noting that the alpha blending is a bottleneck, but it seems to run orthogonal to the algorithmic changes we're discussing. I am thinking both of these solutions could benefit. Perhaps we will find one-hundred ways not to do it, and then the perfect solution will fall into place. :) I will explore TkpPutRGBAImage after ruling out the fix for the changing widgets. If we can get 40% performance increase with a correct fix that is a nice place to start. TkpPutRGBAImage will impact much broader surface area, but I am happy to explore that direction. oehhar added on 2026-06-03 09:03:35: Marc, thanks for the post. This is also my knowledge on the MS-Windows side. In addition, a general solution, and not only for scrollbars, would be great. Thanks, Harald marc_culler (claiming to be Marc Culler) added on 2026-06-02 19:16:45: One thing about chatbots is that they **love** to chat with you. An endless
re-design cycle with ever-expanding scope is just dandy with them. I suggest
that we give the chatbots a rest and think about the core problem with brains,
looking for a pleasing compromise between simplicity and improvement.
Here are a few ideas.
Claude claims "this produces 6,000-12,000 GDI create/destroy cycles per second".
Maybe we could start there. I have no information about how expensive these
so-called cycles are, but let's assume that is the key issue.
Looking at TkPutImage in the windows port I see:
bitmap = (HBITMAP)SelectObject(dcMem, bitmap);
BitBlt(dc, dest_x, dest_y, (int) width, (int) height, dcMem, src_x, src_y,
SRCCOPY);
DeleteObject(SelectObject(dcMem, bitmap));
DeleteDC(dcMem);
TkWinReleaseDrawableDC(d, dc, &state);
where the bitmap which appears as an argument to SelectObject was created by
CreateDIBitmap. I guess Claude is referring to CreateDIBitmap and
DestroyObject as the cycle.
If CreateBitmap and DestroyObject are what is slow, why not cache
the bitmap in TkPutImage, so if TkPutImage were called many times with the
identical image it could avoid all but one of the CreateBitmap and DeleteObject
calls?
A second idea would be: since Schelte finds that the bottleneck is the alpha
blending, rather than the creation and destruction of bitmaps, how about
trying to implement TkpPutRGBAImage for Windows, so that the GPU could be
used to do the alpha blending. I am sure Windows supports using the GPU
for such a task.
A third idea would be to treat the special case of an image which is 1 pixel
high (in the vertical case), or 1 pixel wide (in the horizontal case), as the
primary application for the tiling construction, and try to optimize for that
case. I believe that using StretchBlt might allow tiling an entire scrollbar
trough with one function call instead of 6000 -12000 calls to BitBlt.
There are probably many other, possibly relatively simple, optimizations which
humans could come up with, if they were to think about it.
mtmcp_ added on 2026-06-02 14:14:49: Working with Claude Opus 4.8 and the review notes provided another issue surfaced. I will show the summary info here so folks can get a sense of the logic. The issues
Solution space Four directions, roughly in increasing ambition. They compose. A. Opacity gate (correctness floor, cheapest). Determine once whether the element renders fully opaque over dst (you can derive it from the photo instance's COMPLEX_ALPHA/alpha-table that already exists, or by scanning the source pixels actually used at build time). Only use the cross-frame composite cache when opaque — and when opaque, drop the seed copy. For non-opaque elements, don't keep a cross-frame composite at all. This alone makes the cache correct; it just doesn't accelerate translucent elements. B. Decouple the cache from the background = "tile once, blend per frame." The GDI thrash comes from ~200 per-tile DIB create/blit/destroy cycles, not from compositing. So cache the tiled element as an RGBA image (one expensive tiling pass, cached by geometry+state), then each frame do a single BlendComplexAlpha + TkPutImage of that whole RGBA buffer onto the live d. You re-read the live background every frame (so it's always correct, cases #1–#3 all handled), but you've still collapsed 200 ops → ~1. This is the correct-and-fast answer for translucent elements and pairs naturally with A (opaque → XCopyArea fast path; translucent → cached-RGBA + one blend). C. Hierarchy/graph + bounds-intersection invalidation (the approach you gestured at). Within a DrawWidget pass, track the bounding boxes of elements that cache-missed (i.e., changed) this frame. Before letting element E take its fast path, if any already-drawn element's changed-bbox intersects E.dst, force a miss for E. This is pure rectangle intersection — no pixel readback — and, crucially, it is complete for the in-widget case: everything in d came from an element draw this pass, so "did the background under E change?" is exactly "did any lower element re-render under E's bounds?" This cleanly kills case #1. It needs a small amount of per-pass scratch state (a list of dirty rects), which means threading a context through the draw, i.e., the "re-think in terms of the graph" you mentioned. D. Fix the cache scope (structural, addresses case #2). Move the cache off the shared ImageData. Minimum viable: add window identity to the key (and don't rely on the per-frame d, which is freshly allocated each frame — key on Tk_WindowId, which is stable). Better: a small per-window cache (Ttk elements are intentionally stateless, so this means a side table keyed by (element, window) hung off WidgetCore, or a tiny N-entry LRU on ImageData). Without this, even a correct per-frame algorithm can cross-contaminate two widgets sharing a style. Claude's recommendation: A + D as the correctness baseline (opacity gate + window-scoped key — small, shippable, makes the current code correct), then B for translucent performance and/or C if you want translucent elements to keep a cross-frame cache. A+D is a modest delta on the existing patch; B and C are the "re-think." Notably, A+B together let you delete the seed-copy hack entirely, which removes the whole baked-background failure mode by construction. marc_culler (claiming to be Marc Culler) added on 2026-06-02 14:11:30: Harald, > What I understood from the patch is that (semi-)transparent pixels > are always slow, as they are painted pixel by pixel. That is not correct. In generic/tkImgPhInstance.c there is conditional code bracketed by #ifndef TK_CAN_RENDER_RGBA ... #endif which is responsible for blending an image with transparency over an opaque image. When TK_CAN_RENDER_RGBA is not defined the function BlendComplexAlpha is called. It does the compositing operation on a pixmap in CPU memory. This blending involves copying a rectangle from the window into a pixmap, blending the pixmap with the image, using the image's alpha channel, and then copying the pixmap back to the window. This all happens in the CPU. When TK_CAN_RENDER_RGBA is defined, the function TkpPutRGBAImage is used. That is a platform-specific function which can use the graphics operations provided by the OS to do the blending. The only platform which defines TK_CAN_RENDER_RGBA at the moment is macOS. (Wayland will also use it, as the Wayland port has to interact directly with the Mesa graphics driver using OpenGL.) The macOS port draws directly to the window's backing store CGImage, so it can use Apple's compositing features to do the blending, and it happens in the GPU. mtmcp_ added on 2026-06-02 13:38:57: Thanks everyone for testing this and providing much needed feedback. Several ideas have surfaced which need more rigorous testing.
oehhar added on 2026-06-02 11:39:27: What I understood from the patch is that (semi-)transparent pixels are always slow, as they are painted pixel by pixel. The patch reduces the number of semitransparent pixels by cacheing the image, if the image is composed of multiple images which are superposed and have semi-transparent pixels. If there are any pixels not covered by the cached image (and a background shines through) the performance is bad (but less bad than without the patch, if any semi-transparent pixels change status to non-transparent par the overlay procedure). The fact that semi-transparent pixels are slow (on MS-Windows) is not addressed. Only, the number of those pixels is reduced. Does this make sense? Thanks for all, Harald sbron added on 2026-06-02 10:10:33: I ran the diagonal sweep test with both Tk 9.0.0 and the tip-754-ttk-composed-pixmaps-cache branch (commit d2bcc173) on linux x11. There doesn't seem to be a significant change in the perf numbers reported. Resizing is horribly slow in both cases, compared to a modified version of the script that doesn't use semi-transparent pixels. In that last case, the branch version is more than twice as fast as 9.0.0 and more than 20 times as fast as the original script. Original script:
Modified script:
This is the diff for the modification I used:
marc_culler (claiming to be Marc Culler) added on 2026-05-31 21:55:23: I tested the bugfix branch on Windows with a high-dpi EVO display scaled to 200%. The test script ran and the scrollbar slider was (barely) visible as it is on other platforms. So I take that as empirical proof that there is not problem with XCopyArea on Windows with a high-dpi display, although I can't explain what is making it work. marc_culler (claiming to be Marc Culler) added on 2026-05-31 21:07:46: I guess there is currently no issue with high-dpi screens in linux, because the only scaling possible is to change the screen resolution. For Tk a pixel in a pixmap is equivalent to a pixel on the screen, and Xorg manages that via screen resolution. A default 200x200 Wish window is not large enough for the title bar to hold the default window controls on a high-dpi screen, but that is true no matter how you change resolution. So XCopyArea should work fine in all cases. marc_culler (claiming to be Marc Culler) added on 2026-05-31 15:17:14: Hi Mason, I also worked on the code, after your patch, and added comments to try to explain what was going on there after doing my best to understand the situation. I will commit my changes and you can look at them. However, I think there are a couple of serious issues with this change which I do not think have been discussed. 1. The code assumes that the destination rectangle will look the same whenever the widget containing the image element has the same state. I don't think that is always a valid assumption. I think it might be valid it the widget is opaque, so the background of the destination rectangle gets completely drawn when drawing the parts of the widget that are drawn before the destination rectangle is tiled. But I can imagine cases where that is not the case. For example, a scrollbar with a transparent trough, through which one can see the window content underneath the scrollbar. That content might change without any change to the scrollbar, for example if the content was animated. And in that case the cached pixmap would not match its surroundings if copied into the widget. 2. It is not clear to me that the patch handles high-dpi screens correctly. When you create an mxn pixmap it has m*n pixels, each represented by 4 bytes of CPU memory. But an mxn rectangle on a 2x high-dpi screen has (2*m)*(2*n) screen pixels, representing m*n logical pixels, stored in GPU memory. I believe that TkPutImage, because it is using image drawing code provided by the OS, can render the mxn pixmap into a (2*m)x(2*n) framebuffer rectangle by linearly scaling the pixmap image. I could certainly be wrong, but I suspect that XCopyArea does not handle the scaling the same way. (I do know that this issue is the key blocker for extending XCopyArea so that it can work with pixmap source or destination drawables.) What do you know about #2? How is the high-dpi rendering handled by the patched code versus the previous code? I don't know anything about BitBlt (used by TkPutImage on Windows) but Gemini told me that: "The operation is ultimately translated into a GPU instruction or a high-speed contiguous chunk transfer (such as a hardware-accelerated memcpy) via the Direct3D/DXGI stack or the Desktop Window Manager". That sort of sounds like a Direct3D version of glBlitBuffer, which automatically linearly scales an image if the source and destination rectangles are different sizes. But it apparently can have a CPU source. On the other hand Gemini says that BitBlt does not do scaling and that one would have to use StretchBlt for that. In fact, StretchBlt is used in GdiImage, but it looks like that is only used by the Canvas widget, not by XPutImage. And then there is the question of how the XCopyArea and TkPutImage compare under X11 with a high-dpi screen. So I really don't know where things stand regarding high-dpi screens. While I may be making mountains out of molehills, I would feel much better about this if some of these concerns were addressed by someone who knows more about the situation than I do. oehhar added on 2026-05-31 15:06:17: Mason, I gave you commit rights on the tk repository. Could you update the branch? Thanks for all, Harald mtmcp_ added on 2026-05-31 14:08:48: Marc, I went ahead and made the changes to use The latest changes are in the patch file: mtmcp_ added on 2026-05-31 13:46:06: Marc, I agree re. the blitting comments. I will have them removed. For the I ran a test using the existing script and inferred that it is as it was without the patch because the performance is virtually identical on macOS before/after the patch on macOS. As you stated prior the "improved performance" before was comically because it wasn't doing anything useful at all! However, I am going to do a bit more digging just to make sure we get better coverage. Actually, I am thinking this could be made a unit test if we'd like. I can revise the script for automation it could be a nice regression check. marc_culler (claiming to be Marc Culler) added on 2026-05-30 23:35:43: Never mind -- I see that I can use the performance script to verify that the scrollbar looks like a scrollbar, and assume that it looks the same as you intended. marc_culler (claiming to be Marc Culler) added on 2026-05-30 21:09:07: Mason, Do you have a script that I could run to make sure that your custom scrollbar appears correctly on macOS? marc_culler (claiming to be Marc Culler) added on 2026-05-30 20:59:33: I don't think we need to add a new #define constant which we plan to remove as soon as possible. Why not use TK_NO_DOUBLEBUFFERING which is what avoids essentially the same use of XCopyArea throughout the generic code? It also probably represents what is going on here; that is to say that the fact that macOS draws directly to the screen was probably a big part of why it did not have the big delays in the first place. Also, Claude seems very determined to characterize XCopyArea as a "blit" but I think that is wrong. You cannot blit from a Pixmap (which lives in CPU memory) to the screen. That is always a slow copy operation. A blit moves pixel data efficiently between two framebuffers, both living in GPU memory. I think it is misleading and unhelpful to refer to XCopyArea as a "blit". Could we please remove that comment? mtmcp_ added on 2026-05-30 16:57:52: I have attached the fix so you may test until I can push it (or if you'd like to commit it): ttk-image-macos-fixes.diff It includes a couple other correctness fixes like zeroing variables and using Tcl_Alloc/Free to be consistent. mtmcp_ added on 2026-05-30 16:46:53: Marc, Thank you for finding that issue. I have completed a fix but I don't seem to have write access to the Tk branch. marc_culler (claiming to be Marc Culler) added on 2026-05-29 21:22:30: >> I retry a CI run. If this succeeds, I will merge to main. >> No merge to core-9-0-branch. Harald, Please do not merge to main until the cache code is made conditional on #ifndef NO_TK_DOUBLE_BUFFERING It will break macOS. Please don't do that. marc_culler (claiming to be Marc Culler) added on 2026-05-29 16:08:11: IMPORTANT! I juat remembered why I was so worried about this patch for macOS. It woke me up in the middle of the night. It has a fatal flaw, for macOS only: It will not render the cached Pixmaps! The fact that the test script runs faster is meaningless if the drawing is not happening. The reason that the drawing will hot happen has two parts: (1) The drawing code includes the following line: XCopyArea(Tk_Display(tkwin), imageData->cachedPixmap, d, gc, 0, 0, (unsigned)dst.width, (unsigned)dst.height, dst.x, dst.y); (2) The macOS port of XCopyArea is currently incomplete and does not support a source or destination drawable which is a Pixmap. If passed a Pixmap for either it just returns BadDrawable and does no drawing. (See lines 1125-1128 in macosx/TkMacOSXImage.c.) While XCopyArea is used in the generic code to do makeshift double-buffering (copy from window to pixmap, draw on pixmap, copy back to window), that code is conditional, based on the compiler constant TK_NO_DOUBLE_BUFFERING, which is defined by the macOS port only. Finishing XCopyArea for macOS is an important project which needs to be done. But it is non-trivial for a few reasons, a major one being that it needs to deal with Retina displays. I have been thinking about this lately, after it came up in the ticket [4af5ca1921]. I think the simplest approach would be to require that pixmaps are always using logical pixels and let the OS do the linear scaling to the 2x2 pixels used by a Retina display. That will mean that pixmaps displayed on a Retina display are more pixelated than the origianl and will not be anti-aliased as well. I don't see a practical alternative to that at the moment, but I am not sure it will be adequate for projects like the one discussed in [4af5ca1921]. Another aspect of this is that the scaling which would be required if XCopyArea always treats pixmaps as consisting of logical pixels is that that scaling makes rendering considerably slower. So if the test script were actually rendering the pixmaps, instead of just checking that they are pixmaps and returning BadDrawable when the answer is yes, it might well mean that the macOS speed up would be considerably less, at least on Retina displays. In any case, I think the only reasonable way for this ticket to move forward without breaking the graphics on macOS would be to make the new caching code conditional on ifndef(TK_NO_DOUBLE_BUFFERING), which would exclude macOS, and to revert to the original code when TK_NO_DOUBLE_BUFFERING is defined. I am also curious about whether the original Windows problem is affected by whether the computer has a high-DPI screen. oehhar added on 2026-05-29 09:18:34: Ok, thanks. Now, I understood the tests. You have to resize the window, not scroll the scrollbar. Improvement is directly visible on Windows, thanks for that. I retry a CI run. If this succeeds, I will merge to main. No merge to core-9-0-branch. Any opinions on those proposals? Thanks for all, Harald marc_culler (claiming to be Marc Culler) added on 2026-05-28 18:23:06: On an M3 macbook Air (with Retina display, which affects XCopyArea) running macOS 26.4 I got: Pre-patch: resize-bench (diagonal): 52 steps in 2980.8ms (57.32ms/step) Post-patch: resize-bench (diagonal): 52 steps in 1970.2ms (37.89ms/step) There were no crashes! The non-regression test failures were the usual ones that I see on that system (event-9.18 and event-0.19) plus two new ones (send-13.1 and send-13.2) which I am sure are unrelated. So I would say that the patch looks like a significant improvement for macOS. mtmcp_ added on 2026-05-28 17:38:19: To summarize the performance changes below, I am getting about 40% improvement on Windows/Linux and 20% improvement for macOS. Ideally, the Windows/Linux systems should have sub-100 ms latency on the resizing. Perhaps a follow-up change can be done to improve these systems. However, this minimally-invasive change seems to provide improvements to performance across the board. mtmcp_ added on 2026-05-28 17:17:32: Some macOS performance stats also show an improvement in performance after applying the patch:
mtmcp_ added on 2026-05-28 16:15:16: I have updated the performance test script so that it captures performance to quantify the results of the change. Here is my local results on Windows and Linux (WSL2):
oehhar added on 2026-05-28 09:46:43: The comments so far: Schelte:I could not detect any difference in my testing. But that also means I found no problems with the change. Mason:When running the performance test script linked at the bottom of TIP 754, in order to see the lag without the patch, one must repeatedly drag the diagonal grip to resize the window. On Windows, the perceptible lag is quite noticeable without the patch. However, ideally this should be quantified so I can update the test script to provide performance metrics. Marc:What bothers me is that the problem seems to only be for Windows but the solution affects other platforms and I don't think that it has been carefully tested on, say, macOS, to see whether it improves or degrades performance. The macOS platform does not use any pixmaps when drawing to a window, and I think that may be an important difference. | |||||||||||||
Attachments:
- test_scrollbar_perf.tcl [download] added by mtmcp_ on 2026-06-13 11:35:19. [details]
- test_sections.patch [download] added by erikleunissen on 2026-06-11 12:29:01. [details]
- tk-9.0.3-ttk-image-el-cache.diff [download] added by mtmcp_ on 2026-06-09 16:07:50. [details]
- ttk-layout-node-cache.diff [download] added by mtmcp_ on 2026-06-09 14:19:03. [details]
