71ad7bf762
The ONNX session was built with intra_threads(1) unconditionally. That's correct for DirectML (compute is on the GPU), but on the CPU execution provider it pinned all matmul/conv work to a single core — ~1.78s/image, ~14s for an 8-image batch. Derive the intra-op thread count from available_parallelism() for the CPU EP (leaving 2 logical cores free for the UI and a possible concurrent scan; tagging is the lowest-priority worker, so heavier workers are idle when it runs). DirectML/Auto keep a single thread. Measured ~2.7x on a representative machine (14s -> 5.3s per 8-image batch); sublinear because swinv2 inference is memory-bandwidth-bound. The selected thread count is logged at load.