AI model speed is primarily measured by Time to First Token (TTFT), which indicates initial response speed; Tokens Per Second (TPS), reflecting sustained generation rate; and End-to-End Latency (E2EL), which is the total time for a complete response.