Fair point, for sure.. no simple answer IMO
I guess this depends on use case.
RAG works well for some apps, but for other fairly deterministic target cases, (eg, customer Q/A, payments, medical, etc.) myself and others have had problems getting good accuracy with RAG vector search over the past year. I.e, the bot will either miss things or make things up (poor recall, poor precision in IR terms), which is well documented by Stanford, Google, et all.
So the question arises.. why throw a bunch of chunks at the LLM if you can give it the whole corpus. At least with context injection, when the LLM isn’t accurate you have less things to tune..
Ok it’s not going to work for vast document / code libraries .. but for small and moderate sized content bases it seems to work great so far… am working on a project that is formally testing this.. ! more soon .. thanks
PS. → to make things even more interesting.. I’ve had good luck with context injection + fine tuning.. and there’s emerging approaches that combine RAG and context injection ! .. Etc etc.
also see:
https://pub.towardsai.net/why-rag-applications-fail-in-production-a-technical-deep-dive-15cc976af52c