Case study 04 / Applied AI

Offline Cassava Advisor

A farming advisor that answers questions about cassava disease, pests and varieties with no internet connection at all, on an eight gigabyte laptop from 2015. The language model inside it is off the shelf and it is the least interesting part. Everything that makes the system work, and makes it safe, is the retrieval layer and the code that catches the model when it invents a chemical.

The Cassava Advisor interface in plain black and green type. A question box reads: My cassava leaves are turning yellow and curling near the top of the plant. What could be causing this and what should I do? Below a green Get advice button, three progress steps, understanding your question, searching the guides and writing your advice, are timed at 1 s, 0 s and 33 s, followed by two guides consulted, IITA cassava disease control at a 66% match and IITA cassava pest control at 60%.
Sources on every answer: one of the four things the researcher found already built when he reviewed it.
Role
Sole engineer, with three field consultations
Timeline
July to August 2026
Stack
llama.cpp, Llama 3.2 1B Q4_K_M, FastAPI, hybrid retrieval
Hardware
2015 Intel MacBook Air, 8 GB, CPU only, fully offline
Context
Africa Deep Tech Challenge 2026, eliminated in round one
Repository
Public

The problem

Nigerian cassava farmers lose yield to diseases and pests that are identifiable from visible symptoms. The expertise to identify them already exists, in published agricultural guides that farmers cannot search, cannot always read, and cannot afford the data to reach.

One constraint shaped every decision: it had to run with no internet, on hardware a rural extension worker might actually own.

That constraint eliminates almost every available approach. No API calls. No GPU. No model above roughly two billion parameters. And at that size the model's own knowledge is unreliable enough to be dangerous, because a fabricated chemical recommendation on a food crop is not a bad answer. It is a harm.

So the project stopped being about the model very early, and became about two questions. How do you get the right passage in front of a small model every time. And what do you do when it ignores the passage anyway.


The decisions

  1. A passage competes on the dimension the question emphasises.

    This is the finding I did not expect and the one I have used since. Asked "what chemical should I spray for mealybug", retrieval returns chemical dense passages regardless of whether they are about mealybug. Topical correctness loses to keyword density.

    The fix was counterintuitive: make the correct passage chemical dense too, by writing a passage that names the wrong chemicals explicitly in order to rule them out. Retrieval score went from 0.640 to 0.856 and it beat the competing fungicide passage. You do not win retrieval by being right. You win it by competing on the axis the question is measured on.

  2. Corpus text has to be written in the shape of the question.

    The highest yield answer began as a bare list of names and figures, and ranked seventy first for "which variety gives the highest yield". Rewritten as a full sentence it moved into the top four. No amount of synonym tuning fixed it, because a list of numbers has almost no semantic structure for an embedding model to match. Only the rewrite did.

  3. Deterministic safety, because prompt instructions do not hold.

    I tried repeatedly to stop chemical hallucination with prompt rules, including placing a critical reminder immediately before the answer with a replacement sentence supplied. The model obeyed the rule and violated it inside the same answer.

    So the real defence is code. Any line naming a chemical that does not appear in the retrieved sources is deleted and replaced with a notice. Prompt rules do work for narrower bans, a rule against Latin names successfully removed a fabricated species attribution. They just cannot be trusted with the thing that actually matters.

  4. Found that safety control failing open, silently.

    The server and the client each derived the same sentence split independently, then matched the results by exact string equality. When an answer truncated without terminal punctuation, which happened routinely at the token limit, the strings never matched and the redaction quietly did nothing. The control appeared to work and did not.

    The fix was to compute character offsets once on the server and send them with the exact string that was checked, so the client derives nothing, and to fail closed on a malformed payload. The principle generalises well beyond this project: never have two components independently derive the same split and then match on the result. Compute once, pass offsets.

  5. Igbo as vocabulary bridging, not as a language model.

    Igbo and English vernacular farmer terms map onto the technical vocabulary in the guides, with diacritics stripped so that "akwukwo" matches "akwụkwọ". Igbo symptom questions retrieve the correct pest and disease passages at healthy scores. Rejected: adding an African language model, because the available ones do not cover Igbo and are not commercially licensed, and full Igbo generation from a one billion parameter model would be unverifiable.

  6. Tested fine tuning rather than assuming.

    A LoRA fine tune trained cleanly: 104 steps, training loss 1.366, mean token accuracy 0.838. Then I evaluated it. It garbled variety yields, contradicted itself within a single answer, invented a crop that does not exist, and answered an out of scope question about goats with fabricated veterinary schedules, which was precisely the refusal that had been in its training data.

    Two hundred and six examples taught it style and trained the safety refusal out of it. I shipped the alternative. Below a few thousand examples, knowledge available at inference time beat knowledge diffused into weights.


What did not work

Diagnosis is unstable at this model size. The same question, with the same retrieved facts, produced four different diagnoses across different models and prompt variants. That limitation is documented plainly in the submitted report rather than smoothed over, because a farming advisor that is confidently wrong is worse than one that admits it does not know.


The outcome

On the target laptop, measured on a freshly restarted machine:

On the target laptop, freshly restarted
Throughput18.06 tokens per second
Peak memory1,446 MB
Base accuracy0.64 acc_norm, ARC Easy

Three field consultations shaped the design: an agriculturist at the Ebonyi State agricultural ministry, a production manager at a cassava ethanol plant in Edo State, and a PhD agricultural researcher who has worked with IITA and collected data from more than four hundred farmers across Nigeria and Ghana.

The researcher sent twenty written recommendations after reviewing the system. Four were already implemented before he saw it: the mechanism for saying it does not know, the offline first architecture, visible sources, and protection against dangerous recommendations. His structured diagnostic format was adopted. The remaining recommendations became the roadmap.

An answer set under four headings. Most likely problem: cassava green mite feeding damage. Why: the yellow chlorotic spots are likely mite damage, which can be confused with cassava mosaic disease. What to check: young leaves, and terminal leaves dropping from the shoot tip. What to do: inspect those leaves and remove any affected. A footer line reads: Advice complete. 4 guide passages used, 34 seconds.
The researcher's structured diagnostic format, adopted after his review rather than designed in from the start.

What the judges said, and what I took from it

The submission was eliminated in round one of the Africa Deep Tech Challenge. It scored 3 out of 10 on originality and did not advance.

The reviewers' reasoning was that the shipped model is a stock Llama 3.2 1B with a text knowledge digest embedded in its chat template and no changes to the weights. They also wrote that the experimentation was strong and unusually rigorous, and noted that the rejected fine tune was documented in full.

They were right about the artifact. Embedding a digest in a chat template is a useful packaging decision and it is not novel machine learning. If the question is whether the model is original, the honest answer is no.

What I got wrong was where I put the work. Everything on this page that I think is genuinely new is in the retrieval layer, the corpus design and the safety architecture. The contest scoring explicitly excluded the retrieval pipeline, the corpus, the Igbo layer and the safety gates, and measured the bare model. I spent two months on the parts the rubric could not see.

That is my error rather than theirs. Read the scoring surface before choosing what to build. If the thing being measured is the model, the originality has to be in the model.


What I would do differently

Started with the field consultations instead of reaching those three people partway through. The single most useful change to the output, the structured four heading diagnostic format, came from the researcher, and it would have shaped the architecture rather than been retrofitted into it.