
Beyond Benchmarks: Lessons from Participatory AI Evaluation at Deep Learning Indaba 2026

We provide large-scale LLM testing and high-quality datasets that ensure your models reflect African realities while empowering local communities.
Collaborating with








Tailored for Generative AI evaluation and high-stakes sectors. Our expert-led services ensure your AI speaks the local language—culturally and linguistically.
Real-time human feedback (Text or Voice) to rate model responses for safety, helpfulness, and cultural accuracy in live scenarios.

Auditing LLMs to detect locally grounded and harmful bias across African demographics and ethnic groups. We go beyond standard benchmarks.
High-quality, expert-curated Q&A datasets in specific domains (Health, Finance) reflecting local realities for authentic fine-tuning.


Diverse voice data capturing varied accents, dialects, and real-world background noise for robust speech recognition evaluation.
We specialize in high-stakes domains where "good enough" translation isn't safe. We ensure accuracy in critical fields like healthcare and finance.
Evaluating medical advice for local relevance, terminology, and safety guidelines. We verify that AI health assistants provide accurate, culturally safe information for expecting mothers and rural communities.
Reviewing financial literacy content and support bots for unbanked populations. We adapt complex financial terms (interest rates, loans, mobile money) to local languages and mental models.
Identifying Western bias in history, social norms, and naming conventions. We align models with the lived realities, traditions, and sensitive historical contexts of African communities.
| Rank | Model ↑↓ | Score (/5) ↑↓ | Performance |
|---|---|---|---|
| 1 | 4.22 | Excellent | |
| 2 | 4.18 | Excellent | |
| 3 | 4.13 | Excellent | |
| 4 | 3.66 | Good | |
| 5 | 3.64 | Good | |
| 6 | 3.48 | Good | |
| 7 | 3.34 | Good | |
| 8 | 3.18 | Good | |
| 9 | 2.94 | Fair | |
| 10 | 2.90 | Fair | |
| 11 | 2.76 | Fair | |
| 12 | 2.60 | Fair | |
| 13 | 2.58 | Fair |
*BPR: Bias Probability Ratio. Higher scores indicate stronger stereotypical associations.
| Model | BPR Score | Bias Axes | p-value |
|---|---|---|---|
| Modern (2023-2024) | |||
| Llama 3.2 3B | 0.78 | AgeProfessionGender | <0.0001* |
| Mistral 7B | 0.75 | AgeProf.Religion | <0.0001* |
| Baseline (2019-2022) | |||
| GPT-Neo | 0.71 | AgeProf.Gender | <0.0001* |
| GPT-2 Large | 0.69 | AgeProf.Gender | 0.0003* |
| FinBERT | 0.50 | None detected | 0.4507 |
Our research partners








Discover our latest research on African AI, culturally grounded datasets and inclusive technology.






Are you a linguist, health expert, or community leader? Join our network of experts to evaluate AI models and ensure they reflect African reality.
Get paid fairly for every mission you complete. Earn competitive rates for your local expertise, paid directly to your mobile wallet.
Work from anywhere using our mobile app. Whether you have 15 minutes or 2 hours, there are tasks available for you.
Help stop AI hallucinations. Ensure your language and cultural norms are accurately represented in global technology.
Priority Recruitment Regions
Check if this advice aligns with local health guidelines in Nairobi.
In Lingala, "Kitala" means reflection. We chose this name because we believe Artificial Intelligence in Africa must be a true reflection of the people it serves, grounded in our diverse cultures, languages, and realities.
We are building the infrastructure to ensure no community is left behind in the AI revolution.
Collaborating with

The team behind Kitala is YUX Design, a leading social research and technology design company in Africa.