Medicine matching
A web tool that takes messy Persian drug names, matches them to the official drug registry, and sends uncertain matches to a person for review.
- Role
- Sole developer
- Type
- Client prototype (freelance)
- Period
- Aug 2026
- Link
- Not public
No public link: this is a client prototype. A demo is available on request.
The problem
Pharmacy and distributor data often contains free-text drug names with mixed Persian/Arabic characters, Persian digits, dosage and form mixed into the name, and typos.
The client needed to map those rows to a source-of-truth registry of about 33,800 drugs (IRC code, brand, generic, manufacturer, prices), with a structured record for each and a way to trust or correct each match.
My role
Sole developer: spec, data model, matching pipeline, API and review UI.
Stack
How the matching works
Learned alias: if a person has corrected this exact normalised string before, reuse that match (confidence 1.0).
IRC code: if the row contains a registry code, match it directly.
Normalise: unify ک/ك, ی/ي, ة/ه, convert Persian digits to ASCII, and collapse spaces.
Extract the dosage form (tablet, capsule, ampoule, syrup…) and dosage (
mg,ml,%…) with keyword lists and regex.Fuzzy match the rest against brand and generic names (Latin and Persian) with trigram similarity, combined into a weighted confidence score.
High-confidence matches are auto-approved. The rest go to a review queue where a person approves, rejects or corrects them. Corrections are saved as aliases, so the system improves with use.
Result
- A working end-to-end prototype with batch CSV upload, single-string matching, a review queue and dashboard stats.
- Tested so far on a small sample of real dirty rows, not yet at production scale.
What I’d do next
Measure accuracy on a labelled sample before tuning the score weights. I’d also compare the rule-based pipeline with an LLM-assisted extraction step for the hardest rows.
