0 alternative definitions
The deployment and execution of generative AI models on edge devices or infrastructure for local inference, reducing reliance on cloud connectivity and resources.