Feel the language before you can hear it.
Undertone is a pronunciation trainer that replaces the ear with the skin. Hold your phone against your neck, murmur a syllable, and the handset vibrates the correction your voice should have made — not a score, not a red cross, and not the same audio played back at you again.
No wearable · No accessory · Runs on the phone you already own
01 The problem
You cannot correct a sound you cannot hear.
From roughly the end of the first year of life, human perception tunes itself to the sound categories of the native language. Contrasts that carry no meaning in that language are progressively filtered out. For an English speaker learning Mandarin, the four lexical tones — where the pitch contour of a syllable changes the word entirely — are precisely such filtered contrasts.
Every mainstream language application answers a mispronunciation with a score, a red cross, or the model answer played back again. Every one of those responses is delivered through the auditory channel that is already failing the learner. The result is predictable: pronunciation is the last skill to develop, the most common cause of plateau, and a primary driver of the frustration that makes adults abandon language study altogether.
It is not a niche problem, and in the United Kingdom it is not a soft one.
- £31.3bn central estimate — research for UK Trade and Investment put up to 3.5% of GDP at stake through export trade lost to language skill deficiencies.
- ~£19bn a year — RAND Europe's modelled uplift to UK exports from removing language barriers with Arabic-, Chinese-, French- and Spanish-speaking partners.
- 18,000+ pupils have passed through the Mandarin Excellence Programme since 2016, and "other modern languages" GCSE entries reached 42,113 in 2025 — a fifth consecutive year of growth.
02 How it works
Three seconds. One murmur.
A correction you can act on.
The loop is ninety seconds a day. It runs on the microphone, processor and haptic actuator already inside the handset — and the most sensitive processing never leaves the device.
Throat-Sense
Hold the phone flat against the side of your neck and murmur. The handset's own microphone picks up phonation conducted through tissue, straight from the larynx — a signal dominated by fundamental frequency and voicing, and largely indifferent to the noise around you, because the path runs through the neck rather than the air.
Signed difference
Your extracted pitch contour is compared against a native reference, and what the system keeps is not a verdict but a signed difference — direction, magnitude and timing. Accuracy is expressed acoustically, as measured tone-production accuracy over time, rather than as a quiz score.
The device vibrates the error
The difference is rendered as a designed vibration. A second tone produced flat instead of rising comes back as an intensity sweep that climbs, so your hand feels the ascent your voice failed to perform. A third tone that fails to dip returns as a pulse that falls and rebounds. You don't feel whether you were right. You feel which way to move, by how much, and when.
Tolerance bands adapt as you improve: wide early, so an approximately correct contour passes without a correction and the signal stays informative; narrower as measured accuracy rises, so the same production now triggers a nudge. The ninety-second loop always addresses your highest-error contrast. A companion music module renders the rhythm, stress and melody of connected speech haptically, operationalising a well-established finding that pitch and musical training transfer measurably to tone-language ability.
The neck path never passes through the room, so the noise that destroys the air signal is simply not on it.
03 Why silence matters
It removes the behavioural barrier, not just the inconvenience.
Because the signal is captured through the neck, a murmur is enough. That is easy to mistake for a convenience feature. It is not one.
Speaking a foreign language aloud is the most anxiety-inducing act in language learning, and foreign language anxiety is among the best-documented negative predictors of achievement in the second-language acquisition literature. Adults study on commutes, in shared flats and in open-plan offices — environments in which loud, repetitive pronunciation drilling is socially impossible.
The result is an inverted allocation of effort: the skill that most requires deliberate practice is the one learners practise least. A trainer that works at murmur volume gives that practice back.
04 What makes it defensible
Six things no competitor currently does.
Corrective haptics, not perceptual haptics
Five decades of sensory-substitution work converts incoming sound into touch so a user can perceive it. Undertone reverses the direction: it takes your outgoing production, computes the signed error, and renders the corrective action. This is the central inventive step.
The phone as laryngeal sensor
A contact-mode signal chain matched to the tissue-conduction transfer function, with automatic contact-state detection, contact-specific gain adaptation, and opportunistic fusion with the air channel where the signal-to-noise ratio permits.
Practice at murmur volume
Silent, private, anywhere. It removes the single largest behavioural barrier to deliberate pronunciation practice, rather than adding a convenience on top of the same loop.
Adaptive tolerance bands
Measured accuracy re-weights the feedback bands and selects the next drill item, so the haptic vocabulary tightens as competence rises. This is what turns a feedback device into a curriculum.
Zero hardware
Every previous attempt to bring haptics into speech training required purpose-built hardware, and that requirement is why none reached consumers at scale. Nothing to buy, ship, charge, lose or fund.
Pronunciation as a measurement
Ordinary practice yields an objective longitudinal record: accuracy over time, most frequent error class, practice consistency. Language teaching currently has no widely deployed instrument that does this.
05 Intellectual property
Patent pending in the United Kingdom.
Two applications have been filed with the UK Intellectual Property Office, covering the two independent inventive steps at the core of the platform. A third is drafted and will be filed within the priority year.
Corrective Vibrotactile Feedback Method for Speech Training by Haptic Encoding of Pitch-Contour Error
Compares a learner's produced fundamental-frequency contour against a reference and renders the signed difference as a parametrically mapped vibration, so that direction, intensity envelope and temporal profile encode the corrective action required rather than a measure of correctness.
Application 1 of 3Contact-Mode Extraction of Laryngeal Fundamental Frequency and Voicing Using an Unmodified Mobile Device Microphone
Captures tissue-conducted phonation with an unmodified consumer smartphone in neck contact, applying a contact-matched signal chain with contact-state detection, gain adaptation and opportunistic fusion with the air-conducted channel.
Application 2 of 3Adaptive Haptic Tolerance-Band Control in a Pronunciation Training System
Covers the closed loop in which measured tone-production accuracy re-weights the feedback tolerance bands and selects the learner's subsequent drill items, so the haptic vocabulary progressively tightens as competence increases.
Application 3 of 3A novelty search was commissioned before filing, not after.
A professional novelty search report was completed prior to filing, covering classifications G09B19/04, G09B21/00, G10L25/90, G10L21/06 and G06F3/01 across Espacenet, WIPO PATENTSCOPE, Google Patents and the UK IPO register, together with a non-patent-literature sweep of the sensory-substitution, vibrotactile-feedback and computer-assisted-pronunciation-training corpora.
The report confirms that the claimed combination — laryngeal pitch extraction from an unmodified handset in neck contact, coupled to corrective haptic difference feedback for second-language tone production — is not disclosed in any granted patent or published application in any jurisdiction. The closest art falls into three families, each distinguishable on the record: tactile aids for hearing-impaired users are perceptual rather than corrective; throat-microphone communications hardware uses a dedicated contact transducer rather than an unmodified handset; and computer-assisted pronunciation training returns scores through screen and audio.
Stated precisely: filed applications establish a priority date and a dated technical disclosure. They do not establish granted rights. Grant in the UK typically follows two to five years after filing, with publication at approximately eighteen months. A novelty search is not a freedom-to-operate opinion, and this page does not present it as one. The commercial moat is designed as overlapping layers — the patent family, the proprietary paired corpus, the institutional contract base — rather than as any single instrument.
06 Validation to date
What has been proven — and what has not.
Each item below states its evidential weight honestly. Confirmed commercial commitment is not blurred into general support, and internal results are labelled as internal.
Values are approximate and are stated as the internal protocol reported them. Not peer-reviewed; the panel was small and independent replication has not yet occurred.
- Two UK patent applications filed with the IPO, plus a completed pre-filing novelty search report, retained.
- 142 screened UK residents surveyed across a five-week fieldwork window, recruited through six distinct community channel families — Reddit, Chinese-Forums, Facebook learner groups, Meetup, Discord and UK institutional networks.
- 15 depth interviews with UK adult learners: nine Mandarin, two Cantonese or Vietnamese, four European-language learners at accent plateau.
- 187 UK waitlist signups from ~1,380 sessions — 13.5% visit-to-signup, against a 6.6% all-industry and 8.4% education median, achieved with no paid acquisition.
- Five signed Expressions of Intent from UK organisations covering all three institutional segments. Classified as qualified pipeline, not as revenue.
- A functional MVP on TestFlight: Throat-Sense capture, four haptic error vocabularies across a forty-syllable Mandarin set, the ninety-second loop, the progress view and the teacher dashboard.
Contact capture holds where air capture fails
Pitch contours from phone-against-neck audio were compared against laboratory reference tracking (Praat autocorrelation on studio air-microphone recordings) across the full Mandarin tone inventory in three acoustic environments: quiet room, café-level noise and public-transport-level noise.
Voiced-frame agreement with the reference sat in the mid-nineties in the quiet condition, remained in the mid-nineties at café level, and stayed above 90% at transport level. The air-microphone comparator degraded sharply over the same range, falling into the low seventies as the tracker began mis-tracking on background speech.
Fourteen days, nineteen learners, two independent measures
A fourteen-day cohort study with nineteen UK-based beginner Mandarin learners recruited from the waitlist. Tone-production accuracy was measured pre and post on a held-out syllable set, scored by two native-speaker raters blinded to condition and timepoint, and independently by acoustic contour distance from native targets.
Blinded rating showed substantial improvement across the cohort, with the largest gains on the second–third tone contrast — the error class the haptic vocabularies were designed around. Inter-rater agreement was substantial on the conventional interpretation of kappa.
Limitations, stated rather than buried: nineteen participants, fourteen days, no control group, self-selected from a waitlist, run by the company itself. These are internal engineering and learning results, not peer-reviewed, and independent replication has not yet occurred. A controlled efficacy study with a UK university phonetics or applied linguistics department is a funded deliverable, budgeted at £2,500 in Year 1, £3,500 in Year 2 and £5,000 in Year 3.
07 For schools, universities and employers
The first objective instrument for pronunciation.
Language teaching currently assesses pronunciation impressionistically. Teachers judge by ear under time pressure; examination boards do the same; institutions hold no longitudinal record of any individual student's prosodic development.
Undertone's dashboard produces one automatically, as a by-product of ordinary practice: per-student, per-contrast tone accuracy over time, most frequent error class, and practice consistency. It needs no equipment beyond the phones students already carry — which means deployment is a licence, not a procurement of hardware.
More than 60% of state secondary schools report modern foreign languages recruitment difficulty. A stretched department is precisely the one most receptive to a tool that automates a measurement currently done by ear.
dip-failure 78%
rising-deficit 64%
flat-contour 91%
rising-deficit 52%
falling-deficit 83%
Illustrative view. Measured tone-production accuracy and most frequent error class, generated automatically from ordinary practice — no equipment beyond the phones pupils already carry.
| Institutional licence | Benchmark against the UK market | Price |
|---|---|---|
| Schools per-pupil licenceFull practice loop, class grouping, per-contrast reporting | UK secondary single-subject per-pupil licences observed at roughly £2–£11 per pupil per year | £12 / pupil / yr |
| Teacher & institution dashboardSite licence, unlimited staff seats, longitudinal export | Per-school unlimited-user models observed at roughly £550–£1,200 per subject per year | £950 / school / yr |
| Corporate accent-intelligibility cohortStructured twelve-week programme, measured outcomes | UK specialist studio accent programmes published at £1,350–£5,995 for five to fifteen sessions | £249 / employee |
Institutional sales run pilot-first, on the procurement reality the sector actually operates under: Department for Education research finds 56% of technology investment decisions are made at school level, through a process that includes an explicit piloting and trialling stage.
08 Pricing
Three buyer groups. One technical core.
The same engineering that produces contact-mode pitch extraction serves the consumer subscription, the schools licence, the corporate programme and the SDK simultaneously. Because there is no hardware, margin is pure software margin.
Free
A full daily practice loop and a limited sound set, supported by sponsored practice-break placements. A free user is inventory and word-of-mouth, not a failed conversion.
- Ninety-second daily loopIncluded
- Throat-Sense captureIncluded
- Limited sound setIncluded
- Sponsored placements£24 / '000
£9.99 / month
Ad-free, full contrast inventory, adaptive tolerance bands and the complete progress record. Priced inside the £9–£17 UK corridor for a serious pronunciation app.
- Everything in Free, ad-free—
- Full tone & contrast inventoryIncluded
- Language & contrast packs£14.99
- HSK exam sprint module£29.99
From £950
Schools, universities and employers licence the measurement. Speech-technology firms and phonetics departments licence the engine and the corpus.
- Schools per-pupil licence£12 / pupil / yr
- Dashboard site licence£950 / yr
- Corporate cohort£249 / employee
- Haptic engine SDK£1,500 / month
- Anonymised corpus licence£7,500 / yr
Nine streams across three buyer groups. Gross margin modelled at 91.8% in Year 1, 90.4% in Year 2 and 91.0% in Year 3 after platform commission and merchant fees. Prices shown are UK pricing and are indicative pending launch.
09 Trajectory
Each phase funded by the one before it.
Year 1 is funded by £65,000 of committed founder capital and by revenue commencing in Month 7. Years 2 to 5 are funded from reinvested operating profit, with the founder taking no salary for the first three years.
Build, in-house
A £24,000 build budget across eight overlapping phases: haptic pattern design and perceptual validation, native-speaker reference recording and phonetic annotation for the forty-syllable Mandarin set, security review and penetration testing, and the teacher dashboard front end. Everything that constitutes protected IP is built in-house.
Launch, retention, first pilots
Revenue commences. Focus moves to retention mechanics, funnel conversion, the first institutional pilots, and an HSK exam sprint module timed to examination sittings. R&D runs at 28% of operating expenditure.
Tone inventories, music module, SDK
Cantonese, Vietnamese and Thai tone inventories. The music and rhythm module at full scope. SDK packaging that makes the C++ core licensable to third parties. The controlled efficacy study with a university partner. R&D rises to 33%.
Non-tonal contrast packs, Ireland
Four precision contrast packs — Spanish trilled R, French nasal vowels, Arabic emphatics, Japanese vowel length. International expansion begins with Ireland. R&D at 38%, headcount at eleven.
Germany, the Netherlands, the US and Australia
Continued expansion, with the intended outcome a strategic acquisition in the five-to-seven-year window — valued on the patent family, the proprietary paired corpus and the institutional contract base rather than on consumer subscriber count alone.
Four years teaching English to adults in London. The fourth line on this list is where the problem was found.
10 The founder
I could hear that my tones were wrong when a native speaker told me so. I could not hear how they were wrong — and every application answered a failed attempt by playing the correct version again, as though the problem were that I had not listened carefully enough.
Daliia Gaskarova speaks Russian and Tatar natively, English fluently, and Chinese at an elementary level. She spent four years teaching English to adults in London, watching the same wall from the other side: Russian speakers who could not produce the English th, Arabic speakers who could not separate p from b. Learning English is the reason she was able to move to London, study, and build a career here — and she is acutely aware that the people around her who did not acquire it are living different lives for no other reason.
11 Early access
Join 187 UK learners on the waitlist.
The MVP is on TestFlight and beta places are allocated from the waitlist. Tell us the language you're working on and we'll match you to the right cohort.
Raw voice audio is processed on your own device. Undertone is an educational training platform, not a clinical instrument: it measures pronunciation accuracy for the purpose of language learning, and does not diagnose, screen, treat or monitor speech, hearing or developmental conditions.