TTFT, prefill vs decode, the KV-cache memory bill, quantization, multi-LoRA serving, and how vLLM, SGLang, llama.cpp and MLX run them across NVIDIA, AMD and Apple silicon.
Our IJETS 2024 paper — combining LLaVA-Med with Retrieval Augmented Generation to build a clinically grounded visual question answering assistant for CT and MRI imaging.