I cut a RAG pipeline’s response time from 90 seconds to 4. Never touched the model
Last year I worked with an AI startup, an Oxford spinout. Their product answered research questions through a RAG pipeline. It worked, but every query took around 90 seconds. Long enough that users were bailing before the answer even loaded. The obviou…