29 September 20268 min readJeevitha Nanepalli

My agent kept pitching ₹1,480 fungicide until Hindsight remembered "no"

How we built KhetSmriti, a field officer's agent on Hindsight agent memory: per-tag observations, date-bounded recall and a Memory OFF switch that shows what memory adds.

On 8 April, Ramesh Goud turned down a ₹1,480 bottle of Amistar Top for the leaf blight on his chilli. "Too expensive, I'll wait and see." Ten days later the blight had spread to 40% of his plants and he was looking at a 20% smaller harvest. On 21 April the same field officer came back with a ₹420 copper fungicide and a neighbour's yield result, and Ramesh said yes.

Then a new officer inherits the territory in September. What do they pitch? Without memory, the premium fungicide again.

We built KhetSmriti ("field memory") so that never happens. It's an agent for field officers at agri-input distributors in Telangana, and the entire product is its memory, which runs on Hindsight, Vectorize's open-source agent memory system.

Same farmer, same model, one flag: the brief with Memory OFF vs Memory ON
Same farmer, same model, one flag: the brief with Memory OFF vs Memory ON

What the system does

A field officer covers dozens of farmers across several villages. The loop has three steps:

  1. Before a visit, the officer opens a farmer's page and gets a pre-visit brief: what to check, which products to recommend, how to pitch them, and a list of claims, each tagged with the date of the visit it came from.
  2. After the visit, the officer records a 30-second voice note in Telugu, English or a mix. Groq's Whisper transcribes it, an LLM structures it (crop, stage, issue, products, objection, reaction), the officer corrects it, and it's retained in Hindsight.
  3. Weeks later, the officer records the crop outcome. That's retained too. That's the part that turns "advice given" into "advice that worked".

The stack is Next.js, Groq (openai/gpt-oss-120b with a qwen fallback), zod everywhere, and Hindsight Cloud. One rule held the codebase together: every Hindsight call lives in src/lib/memory.ts, and every call is timed and pushed to a ring buffer that the UI shows in a live Memory panel. If you can't see the memory working, you don't trust it.

KhetSmriti architecture: where Hindsight sits
KhetSmriti architecture: where Hindsight sits

The real design problem: memory that adds up

The obvious design is to retain every visit and recall "stuff about this farmer". That works for one farmer. It falls apart the moment you ask the question that actually matters to a distributor: what works for chilli blight in Chevella?

Three decisions make it work.

1. Tag every memory four ways

Every visit is one Hindsight document, tagged by farmer, village, crop and officer, plus a kind tag that separates advice from results:

// src/lib/memory.ts
export function visitMemoryItem(visit, farmer, village, catalogue): MemoryItemInput {
  return {
    content: buildVisitNarrative(visit, farmer, village, catalogue),
    document_id: visitDocumentId(visit.id),    // "visit:V002": re-seeding replaces, never duplicates
    timestamp: isoTimestamp(visit.date),       // the real visit date, so temporal recall works
    context: "field visit",
    metadata: memoryMetadata(visit, village),  // { farmerId, visitId, village, crop, officerId }
    tags: visitTags(visit, village),           // farmer:, village:, crop:, officer:, kind:
    observation_scopes: "per_tag",
  };
}

Two details matter more than they look. document_id is deterministic (visit:V002, outcome:V002), so re-running the seed script replaces documents instead of duplicating them. I could re-seed as often as I liked while tuning the narratives without ever wiping the bank. And timestamp is the visit date, not the time I called retain. Without it, six months of history would all look like it happened this afternoon.

The content itself is plain English, not JSON. I build a narrative from the structured visit: "Field visit on 2026-04-08 to Ramesh Goud (F001) in Chevella… Advised Amistar Top (P006, premium tier, ₹1,480 per 500 ml…). Farmer objected: price too high…" Hindsight extracts facts from text, so I give it text that a person would write.

2. per_tag observation scopes

This was the single most useful setting in the project. Hindsight consolidates retained memories into observations, which are higher-level beliefs that grow over time. With the default scope, an observation is tied to the exact tag combination of the memory. Since every visit has a unique combination (different farmer, or crop, or officer), observations would never accumulate across visits.

With observation_scopes: "per_tag", each memory feeds a separate observation for each of its tags. So Hindsight builds, independently:

  • a picture of Ramesh (farmer:F001): rejects premium on price, accepts when shown a neighbour's result;
  • a picture of Chevella (village:chevella): copper fungicide controls chilli blight;
  • a picture of chilli across villages.

That's what makes the village insights screen possible. It calls reflect with a village tag and asks three fixed questions: what works for the top problem, how to handle objections, and what to watch next month. On my seed data, reflect found that pink bollworm returns to Moinabad cotton every August, and that Chevella farmers accept advice when it comes with a neighbour's yield result. I never wrote either rule anywhere. They were just in the history.

3. Two recalls, and no remembering the future

A brief needs two kinds of context: this farmer's history, and what happened to other farmers growing the same crop in the same village. So buildBrief runs them in parallel:

// src/lib/brief.ts
const [farmerMemories, ...similarByCrop] = await Promise.all([
  recallFarmer(farmer.id, query, visitDate),
  ...farmer.crops.map((crop) =>
    recallSimilar(village.slug, crop, `${crop} problems in ${village.name}: which advice
      and products worked, which failed or were rejected on price?`, visitDate),
  ),
]);

recallFarmer uses farmer:<id> with any_strict. recallSimilar uses village:<slug> + crop:<slug> with all_strict, which means "tagged with both, and never untagged memories". Then I drop anything from the similar set that belongs to the same farmer, because his own history is already in the first list.

The subtle problem is time. My demo replays Ramesh's history visit by visit, and a brief as of 8 April must not know how July turned out. Plain recall has no idea what "now" is, so it would happily surface a July outcome. The fix is two lines: pass the visit date to Hindsight as queryTimestamp for recency scoring, then filter out any hit whose occurredAt is after the visit date:

...(asOf ? { queryTimestamp: `${asOf}T23:59:59+05:30` } : {}),
// …
return asOf ? hits.filter((h) => !h.occurredAt || h.occurredAt.slice(0, 10) <= asOf) : hits;

This only works because every retain carried the real visit timestamp. Temporal correctness has to be designed in at write time; you can't bolt it on at read time.

Before and after: same farmer, same model, one flag

The brief page has a Memory toggle, and /demo shows both briefs side by side. They call the same endpoint; the only difference is memoryEnabled. When it's off, buildBrief skips both recalls entirely and tells the model there's no history. Nothing is faked.

For Ramesh, on 28 September:

Memory OFF: "No history with this farmer yet." Four generic budget products for chilli and cotton, a zinc supplement "for likely deficiency", and confidence: low. Every product's evidence reads "Standard recommendation for these crops".

Memory ON: 47 memories about Ramesh and 57 similar Chevella cases recalled. Two products, not four: Blitox for chilli blight ("controlled after two sprays at 500 g/acre", cited to 2 May) and Ulala for cotton ("hopper burn stopped, squares retained", cited to 14 August). The pitch leads with the ₹420 price and the neighbour's result, because that's what worked on 21 April. And it lists Amistar Top's rejection and the 40% spread as a claim, cited to 18 April, so the officer knows what not to say.

The first brief is a reasonable chatbot. The second one is a colleague who's been to this farm five times.

Guardrails, because it's advice about pesticides

An agent that recommends crop chemicals can't be allowed to improvise. Hindsight helps at the bank level: the bank has a mission ("I am the field memory of an agri-input distributor…") and four directives that it enforces during reflect: only recommend catalogue products, never state a dosage that isn't in the catalogue, always cite the visit date behind a claim, and say plainly when there's no history. npm run setup:bank syncs them idempotently.

I don't rely on that alone. After every LLM call:

// src/lib/guardrails.ts
export function enforceCitations(brief: Brief, knownDates: ReadonlySet<string>): ClaimGuardResult {
  const unverifiedClaims: string[] = [];
  const claims = brief.claims.map((c) => {
    if (c.sourceVisitDate && !knownDates.has(c.sourceVisitDate)) {
      unverifiedClaims.push(c.text);
      return { ...c, sourceVisitDate: null };
    }
    return c;
  });
  return { brief: { ...brief, claims }, unverifiedClaims };
}

knownDates comes from the recalled memories. If the model cites a date that isn't in what Hindsight returned, the date is stripped and a warning is shown. Unknown product IDs are dropped the same way. Every LLM response goes through zod; on failure, it retries once with the validation error appended, then falls back to the second model, then returns a typed error. The Memory ON sample above took two attempts. The first output failed the schema, and the officer never saw it.

Lessons

1. Tags are your schema. Design them before you retain anything. Retrofitting tags means re-seeding the whole bank. I settled on farmer:, village:, crop:, officer:, kind: on day one and every query I needed later fell out of them.

2. Use per_tag observation scopes when memories belong to several entities at once. A visit is about a farmer and a village and a crop. Let each of them learn from it.

3. Timestamp at write time with the real event date. It's the difference between "recent memories" and "correct memories", and it's what makes as-of queries possible.

4. Build the Memory OFF switch. It keeps you honest. Any extra context you sneak into the prompt shows up in the "before" too, so the side-by-side only shows what memory actually adds.

5. Guardrails only cover the fields you check. This is the one I'm still fixing. My citation check covers the structured claims array. In the Memory ON sample, the free-text pitch says "see 2024-05-02 result", a year the model invented, and nothing caught it because the pitch isn't a checked field. The next version scans free text for dates too.

The code is on GitHub. If you're building an agent that needs to remember people over months, start with the Hindsight docs and Vectorize's explainer on what agent memory is and why stateless agents fall short. Then work out your tags before you write anything else.

JJISPL (JJ Infotech Solutions Pvt Ltd) is a Hyderabad-based technology company building digital solutions for agriculture and allied sectors.