Inside the Data Engine
September 15, 2026
6 min read

Introducing Nebula: a language platform for Ghanaian institutions

Speech recognition, machine translation, and message routing across Twi, Ewe, and Ga, deployed inside the institution's own perimeter.

AdwumaTech AI
Editorial Team
AdwumaTech AI brand mark beside a luminous global AI network in ivory and gold.

Today we are introducing Nebula, AdwumaTech's platform for running custom AI applications in Ghanaian languages. Nebula provides speech recognition, machine translation, and message routing across Twi, Ewe, and Ga. It powers live interaction channels on WhatsApp and voice, engaging a customer in the language they speak and writing a structured English record to the systems behind the channel.

Nebula is trained on speech collected in Ghana. Part of that corpus is public. mGhana-ST is AdwumaTech's own native-speaker collection across Twi, Ewe, and Ga, released under MIT license on Hugging Face, and named Best African Dataset at the Deep Learning Indaba 2026. UGSpeechData, compiled by University of Ghana researchers across five Ghanaian languages, was restructured and published by AdwumaTech for the research community. The remainder of the training corpus is proprietary.

Nebula deploys entirely inside the client's security perimeter, in a private cloud or on local hardware. AdwumaTech holds ISO 27001 certification and aligns to ISO/IEC 42001 and GDPR.

Snapshot

PlatformNebula
LanguagesTwi, Ewe, Ga
CapabilitiesSpeech recognition, machine translation, message routing
ChannelsWhatsApp, voice
Record outputStructured English, written to the client's backend systems
DeploymentClient-controlled private cloud or local hardware
Open data componentsmGhana-ST, UGSpeechData (MIT, Hugging Face)
GovernanceISO 27001 certified; ISO/IEC 42001 aligned; GDPR aligned

A training corpus collected in Ghana, published where it can be inspected

General-purpose speech systems are trained on corpora in which Ghanaian languages and Ghanaian-accented English are marginal. Accuracy degrades on the speech an institution in Accra or Kumasi actually receives, and degrades further on code-switched utterances, which are the norm in a service call.

Nebula's recognition stack is trained on a corpus assembled for Ghanaian speech. Two components of it are public and carry MIT licenses. mGhana-ST is native-speaker audio in Twi, Ewe, and Ga, aligned to English translations, collected by AdwumaTech engineers in Ghana. UGSpeechData is a University of Ghana corpus spanning five Ghanaian languages, restructured and republished by AdwumaTech. Both are downloadable today and in use by teams with no commercial relationship to AdwumaTech.

The open release sets the standard the rest of the corpus is held to. A buyer evaluating Nebula can read the collection methodology, listen to the audio, and check the transcription conventions before any commercial conversation. Deep Learning Indaba named mGhana-ST Best African Dataset in 2026 on that public record.

One channel, two outputs

A customer speaks Twi into WhatsApp or a phone line. Nebula recognises the utterance, routes it to the correct workflow, and returns a reply in the same language. The institution's backend receives an English record with the fields its systems expect.

The design point is that the language barrier is absorbed at the channel. No downstream system is modified, no staff member transcribes a call by hand, and no record is lost because the conversation happened in a language the CRM does not hold.

Deployment inside the perimeter

Financial services and public sector buyers in Ghana carry data residency obligations that rule out sending customer audio to an external API. Nebula runs in a private cloud under the client's control or on local hardware in the client's facility. Audio, transcripts, and records stay inside the perimeter for the full lifecycle.

AdwumaTech holds ISO 27001 certification, aligns its AI management practice to ISO/IEC 42001, and aligns its data handling to GDPR. Governance documentation is available under NDA for security review.

Where Nebula applies

Nebula fits any workload where an institution serves customers who speak Twi, Ewe, or Ga and keeps its records in English. Five sector patterns follow from that.

Banks and insurers run account enquiry, dispute intake, loan screening, and collections contact through the channel. Every one of those conversations has to end as a structured record in a core banking or CRM system, which is the constraint that rules out a translation layer bolted on after the fact.

Telecommunications volume concentrates in first-line support, bundle changes, and outage reporting, where cost per contact is measured weekly and language handling drives average handling time.

Public sector service delivery covers citizen enquiry lines, benefit and registration intake, and information campaigns that have to reach past the English-speaking population. Residency obligations are strictest in this sector, which makes the in-perimeter deployment model a precondition for procurement.

Utilities carry meter reading submission, billing enquiry, and fault reporting at high volume, today dependent on whichever staff member speaks the caller's language.

Enterprise operations outside these regulated sectors use the same channel for customer support, order and delivery status, field staff reporting from sites where English is the second language, and internal HR enquiry across a distributed workforce.

Getting started

The open datasets are the fastest way to assess the foundation Nebula is built on. mGhana-ST and UGSpeechData are available on Hugging Face under MIT license for any team to download and evaluate.

Institutions scoping a first workload can start with the Deployment Readiness Assessment at adwumatech.ai/deployment-readiness-assessment, which scores a single use case across four pillars and returns a maturity tier before any commitment to build.

Frequently asked questions

What is Nebula?

Nebula is AdwumaTech's language platform for Ghanaian institutions. It provides speech recognition, machine translation, and message routing across Twi, Ewe, and Ga, and runs live customer channels on WhatsApp and voice. Conversations happen in the customer's language and are written to the institution's backend as structured English records.

Which Ghanaian languages does Nebula support?

Nebula supports Twi, Ewe, and Ga for speech recognition, translation, and routing. Ghanaian-accented English and code-switched speech, where a speaker moves between a Ghanaian language and English inside a single utterance, are handled as part of the same pipeline.

What data is Nebula trained on?

Nebula is trained on a speech corpus collected in Ghana. Two components are public under MIT license on Hugging Face: mGhana-ST, AdwumaTech's native-speaker collection in Twi, Ewe, and Ga, named Best African Dataset at the Deep Learning Indaba 2026; and UGSpeechData, a University of Ghana corpus across five Ghanaian languages that AdwumaTech restructured and republished. The remainder of the corpus is proprietary.

Can Nebula run on-premise or inside a private cloud?

Nebula deploys entirely inside the client's security perimeter, either in a client-controlled private cloud or on local hardware in the client's own facility. Audio, transcripts, and derived records remain inside that perimeter for their full lifecycle, which meets data residency requirements that prevent Ghanaian banks and public sector bodies from sending customer audio to an external API. AdwumaTech holds ISO 27001 certification and aligns to ISO/IEC 42001 and GDPR.

What workloads does Nebula handle?

Nebula handles any workload where an institution serves customers in Twi, Ewe, or Ga and keeps records in English. Common examples are account enquiry and dispute intake in banking, first-line support and outage reporting in telecommunications, citizen enquiry and registration intake in the public sector, meter reading and billing enquiry in utilities, and customer support and field staff reporting across enterprise operations.


Datasets

mGhana-ST: https://huggingface.co/datasets/adwumatech-ai/mghana-st

UGSpeechData: https://huggingface.co/datasets/adwumatech-ai/UGSpeechData

AdwumaTech AI
Editorial Team

AdwumaTech AI publishes operational diagnostics, systems research, and implementation insight on enterprise and government AI in Africa and beyond.

Tags

NebulaAfrican Language AISpeech RecognitionMachine TranslationGhana

Explore Our Solutions

Discover how we build high-quality data for frontier AI models.

View our AI solutions