Not Everything Belongs in the RAG

Enterprise AI projects often assume that more data creates a better system. Sometimes the right architecture depends on knowing what should not be indexed, embedded, retrieved, or exposed.

Not Everything Belongs in the RAG

Kraft Foods stands as one of the most memorable projects of my career. It was a competitive enterprise-search project: I was on Team SharePoint, and we were up against Team Google Search Appliance. We won, they lost - big!

But today I am not thinking about MOSS Search, managed metadata, or indexing infrastructure. I am thinking about Kraft’s R&D department. For a brief time, R&D was an obstacle because they refused to let us index their content. Our objective was to pack as much information as possible into the search engine and prove what it could do. It took us a moment to stop and ask: Why?

It turned out that R&D maintained a level of compartmentalized security that would make the Pentagon blush. Even teams within the department were not allowed to know what other teams were developing. Who knew mint-flavored Oreos and Flamin’ Hot Mac & Cheese were so sensitive?

Today, we talk about RAG and vector databases with much the same underlying assumption: everything should be included. Somehow, ingesting the entire institutional corpus is supposed to make the AI experience better. Sometimes “better” means compromised intellectual property, a weakened security posture, and greater risk exposure.

Knowledge architecture is partly about discovery. It is also about preserving boundaries:

Who may know something?
Under what circumstances?
For what purpose?
And with how much context?

Not everything belongs in the RAG.


If you think Pumpkin Spice Mac & Cheese is a must-have flavor, hit that “seem like AI slop” button now


Also published on LinkedIn.