By Pita Ligaiula

Pasifika AI creator and Kanaky Tech founder Kevyn Wahuzue speaks to PACNEWS journalist Pita Ligaiula about the challenges of developing artificial intelligence that respects Pacific cultures, languages, histories and community authority.

In this featured interview, Wahuzue explains why Pasifika AI was created, how it verifies Pacific language information and limits its sources, the risks of AI generating confident but unverified information, and why communities must retain control over their cultural knowledge as AI systems are developed and deployed across the Pacific.

PACNEWS:What motivated you to create Pasifika AI, and what gap did you see in existing AI systems when it comes to Pacific cultures, languages and histories?

WAHUZUE:The gap is not that mainstream systems know nothing about the Pacific. It is that they cannot tell you how much they know, and they do not stop. They produce the same confident register for a well-documented topic and for one where the material simply does not exist in their training data. For a region whose knowledge is unevenly digitised, under-represented in the sources these models learned from, and in many cases deliberately not published, that failure mode is the whole problem. I wanted a tool where the route back to the source stays visible, and where the limits are part of what you see rather than something you discover after quoting it.

PACNEWS: You describe Pasifika AI as being human-centred and culturally conscious. What specific safeguards have you built into the system to prevent it from generating or sharing culturally restricted or inappropriate knowledge?

WAHUZUE:Three, and I would rather describe them accurately than make them sound stronger than they are.

First, the retrieval boundary: when the system searches the web, it consults only a published list of institutional domains, so what it draws on is material those institutions have already chosen to make public. It is not searching the open web, and it is not surfacing what somebody posted in a forum.

Second, vocabulary verification. Every term the assistant uses is tied to the institution it came from, the reference is shown with the answer, and when a word cannot be traced back to a source the assistant says so instead of presenting it as established. Where a dictionary is available, the check goes one step further and becomes mechanical. For example, for Kanak languages, where Kanaky Tech publishes the dictionary, every Kanak word in an answer is compared with a 19,655-entry index and any word that is not in it is flagged to the reader, as described in question 4. That is the template I want to extend, language by language, with the institutions that hold each dictionary.

Third, and most important, an explicit statement of the boundary itself: some knowledge belongs to a family, a clan, a lineage, particular custodians. Being able to ask a model a question does not create permission to receive that knowledge. That is stated publicly in the project statement rather than buried in a technical page.

What I will not claim is that these safeguards are complete. A restriction to institutional sources reduces the risk that the system repeats something it should not have been given; it does not make the system a competent judge of what is restricted. No automated filter can be. That judgement belongs to communities, and it is exactly why the next questions matter more than the technical ones.

PACNEWS:You say the system searches only 59 institutional domains. Why did you choose this approach, and how do you determine which institutions and sources are sufficiently authoritative for inclusion?

WAHUZUE:Because a published, finite list can be argued with. A model that searches everything cannot be held to anything — you cannot inspect it, you cannot contest an inclusion, you cannot notice an absence. A list of 59 institutional domains can be opened by any reader, and someone can tell me it is wrong.

The criteria are deliberately conservative: language academies, universities, museums, archives and regional organisations — bodies that publish under their own name, keep their material accountable, and can be cited. I built the initial list myself, which is a limitation I want stated plainly in your piece: it reflects one person’s judgement about what is authoritative, and it will carry the blind spots of that judgement. Listing an institution also does not mean it endorses the project or has any relationship with me. The list is a sourcing policy, not a partnership claim, and not a guarantee that every generated sentence is correct — readers should still open the references.

PACNEWS:For the nine Kanak languages, you mention a 19,655-entry dictionary index. How does the external language check work, and what happens when the AI produces a word that is not found in the dictionary?

WAHUZUE:For Kanak-language questions the system works against an index of 19,655 dictionary entries across nine languages, drawn from the dictionary we publish. The order matters: dictionary material is retrieved before the answer is generated, not checked afterwards, and Kanak words in the output are then verified against the index.

When a word is not found, the intended behaviour is to flag it as unverified rather than present it as established. That is the whole design — an explicit gap is more useful than an invented answer that sounds complete.

Two limits I insist on. An absent entry does not mean the word does not exist; it means our material does not cover it, and the coverage is uneven between the nine languages. And a match in the index does not establish that a translation fits a particular context. A dictionary is not a speaker.

PACNEWS: How do you deal with differences between Pacific communities over cultural knowledge, language, traditions or historical interpretation?

WAHUZUE:I do not resolve them, and I do not think a system like this should try. Where accounts differ, the honest output is to show the sources and let them differ, rather than average them into a single smooth paragraph — the averaging is where the damage happens, because it quietly erases the fact that there was a disagreement and who held which position. Naming sources is what makes that possible: the reader sees which institution said what, and can go further.

The harder cases are the ones where a difference is not a matter of published record at all, but of standing — who is entitled to tell a particular account. There the tool has nothing useful to contribute, and I would rather it point to references and stop than perform a judgement it has no basis for.

PACNEWS: You have stressed that Pasifika AI has no institutional partner or community mandate. What are the risks of developing Pacific-focused AI without formal community ownership or endorsement, and how could those risks be addressed?

WAHUZUE:This is the question I would most like you to keep in the story, because I have stated the absence publicly rather than waiting to be asked.

The risk is real and it is not hypothetical: a tool that speaks about a people without their mandate can become a de facto reference simply by being available and convenient. Convenience is authority’s cheapest route. Over time, an interface can end up shaping what people outside the region take to be true about it, with nobody having agreed to that.

What I have done so far is partial: restrict the sources, publish the limits, state explicitly that the project does not speak for Pacific communities and that an AI answer does not replace a speaker, a teacher or a knowledge custodian. That reduces the harm. It does not resolve the question.

What would resolve it is governance that I do not currently have — a mandate that can constrain me, including the power to require that something be removed. I would rather say that openly than let the project’s existence imply an endorsement that nobody gave.

PACNEWS: The Pacific AI principles call for AI that is aligned with Pacific values. From your experience building Pasifika AI, what does that mean in practical terms rather than simply as a principle?

WAHUZUE:In practice it has meant accepting worse product metrics.

A system that refuses, that flags an unverified word, that returns five institutional references instead of one confident paragraph, is a less impressive demo. It answers fewer questions completely. Every instinct in this industry pushes the other way, because fluency is what gets measured and shared.

So concretely: it means the refusal is a feature you design and pay for, not an error you apologise for. It means restricting where the system may look even though that shrinks what it can answer. It means publishing the limits on the same page as the capability. And it means separating access from authority — the tool can be free and easy to use without that making it a voice for anybody.

PACNEWS: What do you see as the biggest risk of mainstream AI systems being used in Pacific communities without stronger safeguards around cultural knowledge, language and identity?

WAHUZUE:Not that these systems say something offensive — that is visible and gets corrected. The risk is the quiet one: they answer smoothly about things they have almost no material on, and that answer becomes the default reference for a student, a teacher, a journalist, a public servant. Under-representation in training data does not show up as silence. It shows up as confident invention.

The second-order risk is worse. Invented material gets published, indexed, and read back in by the next generation of models. A fabricated word or a garbled account can become part of the record about a language with very few speakers, and nobody can trace where it entered. Communities with the least digitised material are the most exposed, which is precisely backwards.

PACNEWS: Could Pasifika AI be adapted or expanded through partnerships with Pacific governments, universities, cultural institutions or communities, and what would such a partnership need to look like to ensure communities retain ownership of their knowledge?

WAHUZUE:Yes, and I would welcome it. It is the direction the project has to go if it is to be more than one person’s work.

For me, a partnership that genuinely protected community ownership would need at least these things: the community or institution decides what goes in and what stays out of its own material, and can require removal at any time, without having to argue for it; the dictionary and source material stay under the owners’ control rather than being absorbed into a system they cannot get back out of; the project cannot claim a mandate wider than what was actually given; and there is a way to say no that is not merely advisory.

What I can offer is the working system, the technical build, and a design that is already structured around named sources and stated limits — so a partner’s material stays identifiable as theirs rather than dissolving into a model. What I cannot offer, and would not, is a promise that the technology substitutes for the people who hold the knowledge.

PACNEWS: What would you like Pacific governments and regional organisations to consider as they move from adopting AI principles to actually developing and deploying AI systems?

WAHUZUE:Three things.

a) Ask where a system is allowed to look, and require the answer in public. A vendor who cannot show you their source list is asking you to trust a boundary they will not describe.

b) Ask what the system does when it does not know. Most procurement evaluates the best answer. The behaviour that matters in a cultural or public-service context is the failure case, because that is where the harm lives, and it is almost never tested.

c) And keep ownership separate from access. Deploying a tool is not the same as ceding authority to it. The principles are right that AI should be human-centred and culturally aligned; the practical test of that is whether communities retain the power to say no after the system is live — not only during the consultation before it is bought.