Execution-level optimization of a fixed local AI inference workload on a Google Pixel 7 (Arm-based Tensor G2), studying CPU thread-count.