RAG on Liferay Objects: Building a Distributor-Network Chatbot with a Local LLM
This is LR Tools’ rewrite of a build log by Ankit Srivastava on Liferay.dev, describing a genuinely practical use of Retrieval-Augmented Generation on top of Liferay’s Object framework.
The problem: knowledge locked in spreadsheets
The author’s company manages a global distributor network — different regions, tiers, and sales performance — and answering something as simple as “who’s our top performer in Asia Pacific?” meant digging through spreadsheets or tracking down whoever happened to know. The goal was a chat widget on the company portal that could answer these questions directly, using the actual live distributor data rather than a stale export.
Why RAG instead of fine-tuning
Training a model on the data was ruled out early — it’s slow, costly, and goes stale the moment a new distributor record is added. The alternative, Retrieval-Augmented Generation, keeps the data where it already lives, fetches the relevant records at query time, and hands them to an LLM as context for a plain-English answer. The design constraint that shaped everything downstream: a newly added distributor should be answerable the very next day, with zero code changes.
Step 1: modeling the data as a Liferay Object
The data itself is a custom Liferay Object Definition, Distributor, with 14 fields covering everything a query might touch: Distributor Name, Location, Region, Annual Quota, Last Year Sales, Current Year Sales, Distributor Tier (Bronze/Silver/Gold/Platinum), Established Year, Number of Employees, Active Status, Product Categories, Contact Email, Contact Phone, and Notes. Thirty realistic dummy records were loaded across North America, Europe, Asia Pacific, the Middle East, Latin America, and Africa, with enough variation in tier, size, and performance to make testing meaningful.
Step 2: a chat widget that actually stays current
The widget itself is a Liferay HTML Fragment — a floating chat button that opens a chat panel. The first version hardcoded the distributor data straight into the widget’s JavaScript, which worked but defeated the entire point of RAG: a new distributor added tomorrow wouldn’t show up until someone updated the fragment. The fix was calling Liferay’s Headless REST API at runtime instead:
GET /o/c/distributors?page=1&pageSize=100
On every question, the widget now fetches the latest distributor records from Liferay, converts them into plain-text context, sends that context plus the question to the LLM, and displays the answer. A distributor added five minutes ago is included in the very next query — no redeploy, no cache to clear.
Step 3: moving the model local, for a data-sensitivity reason
Before wiring this up to a cloud LLM API, the author paused on what the data actually was: sales figures, internal quotas, tier classifications, and personal contact details for business partners — none of it public. Sending that to a cloud API means it leaves company infrastructure and gets processed on someone else’s servers, which can run straight into regional regulations, distributor contracts, or internal governance policy, and would reasonably need a legal review before shipping.
The resolution was simpler than negotiating that review: don’t send the data anywhere. Ollama runs open-source LLMs entirely on the local machine — no API key, no cloud account — and exposes a local API server on localhost:11434:
# Allow requests from Liferay running on port 8080
OLLAMA_ORIGINS=* ollama serve
# Pull a capable open-source model
ollama pull llama3.2
With this swap, the same RAG flow applies, but the trust boundary changes entirely: data moves from Liferay to the browser to Ollama, all within localhost, never touching the public internet. Revealing which distributors missed quota, or who’s classified as Platinum tier, exposes commercially sensitive relationships that shouldn’t leave the building without explicit sign-off — running the model locally removes that question altogether rather than answering it.
What it can actually answer
With live data and a local model in place, the widget handles queries like: who exceeded quota last year (ranked by achievement), every Platinum-tier distributor (filtered by tier), which distributors sit in Asia Pacific (filtered by region), the top performer by sales (sorted by last-year sales), which distributors are inactive (filtered by status), and quota-versus-actual comparisons across the whole network (a computed comparison, not a stored field).
The stack, end to end
| Layer | What’s used |
|---|---|
| Data store | Liferay Object (custom Distributor definition) |
| API | Liferay Headless REST (/o/c/distributors) |
| LLM | Ollama, running locally (llama3.2) |
| RAG | Live fetch on every query, plain-text context injection |
| UI | Liferay HTML Fragment (floating chat widget) |
| Auth | Liferay.authToken passed as p_auth + X-CSRF-Token |
This article is LR Tools’ rewrite of the original post by Ankit Srivastava — read it on Liferay.dev for the author’s own framing.
This article is adapted from: Ankit Srivastava, Liferay.dev