0 alternative definitions
A serving optimization technique that dynamically combines incoming inference requests into batches to improve hardware utilization and increase throughput.