Patent pending · United Kingdom

Feel the language before you can hear it.

Undertone is a pronunciation trainer that replaces the ear with the skin. Hold your phone against your neck, murmur a syllable, and the handset vibrates the correction your voice should have made — not a score, not a red cross, and not the same audio played back at you again.

No wearable · No accessory · Runs on the phone you already own

Tone 2 · rising Live
妈 → 麻
−52
Native reference Your production
Accuracy72%
Attempts14
Band±35Hz
Next 马 · tone 3 · dip-failure
Haptic correction · rising-deficit
Your pitch stayed flat. Feel the climb.
Throat-Sense capture · murmur volume
2+1
UK patent applications filed with the IPO, with a third drafted for filing within the priority year
13.5%
Waitlist conversion from ~1,380 sessions — roughly double the education-sector median, on zero paid acquisition
5
Signed Expressions of Intent from UK organisations across all three institutional segments
0
Items of hardware to buy, ship, charge, lose or fund. The phone is the instrument

01 The problem

You cannot correct a sound you cannot hear.

From roughly the end of the first year of life, human perception tunes itself to the sound categories of the native language. Contrasts that carry no meaning in that language are progressively filtered out. For an English speaker learning Mandarin, the four lexical tones — where the pitch contour of a syllable changes the word entirely — are precisely such filtered contrasts.

Every mainstream language application answers a mispronunciation with a score, a red cross, or the model answer played back again. Every one of those responses is delivered through the auditory channel that is already failing the learner. The result is predictable: pronunciation is the last skill to develop, the most common cause of plateau, and a primary driver of the frustration that makes adults abandon language study altogether.

It is not a niche problem, and in the United Kingdom it is not a soft one.

  • £31.3bn central estimate — research for UK Trade and Investment put up to 3.5% of GDP at stake through export trade lost to language skill deficiencies.
  • ~£19bn a year — RAND Europe's modelled uplift to UK exports from removing language barriers with Arabic-, Chinese-, French- and Spanish-speaking partners.
  • 18,000+ pupils have passed through the Mandarin Excellence Programme since 2016, and "other modern languages" GCSE entries reached 42,113 in 2025 — a fifth consecutive year of growth.

02 How it works

Three seconds. One murmur.
A correction you can act on.

The loop is ninety seconds a day. It runs on the microphone, processor and haptic actuator already inside the handset — and the most sensitive processing never leaves the device.

Step 01 — Capture

Throat-Sense

Hold the phone flat against the side of your neck and murmur. The handset's own microphone picks up phonation conducted through tissue, straight from the larynx — a signal dominated by fundamental frequency and voicing, and largely indifferent to the noise around you, because the path runs through the neck rather than the air.

Step 02 — Measure

Signed difference

Your extracted pitch contour is compared against a native reference, and what the system keeps is not a verdict but a signed difference — direction, magnitude and timing. Accuracy is expressed acoustically, as measured tone-production accuracy over time, rather than as a quiz score.

Step 03 — Correct

The device vibrates the error

The difference is rendered as a designed vibration. A second tone produced flat instead of rising comes back as an intensity sweep that climbs, so your hand feels the ascent your voice failed to perform. A third tone that fails to dip returns as a pulse that falls and rebounds. You don't feel whether you were right. You feel which way to move, by how much, and when.

Tolerance bands adapt as you improve: wide early, so an approximately correct contour passes without a correction and the signal stays informative; narrower as measured accuracy rises, so the same production now triggers a nudge. The ninety-second loop always addresses your highest-error contrast. A companion music module renders the rhythm, stress and melody of connected speech haptically, operationalising a well-established finding that pitch and musical training transfer measurably to tone-language ability.

One café. Two signal paths. Recorded simultaneously · illustrative
AIR MICROPHONE PHONATION BURIED THROAT-SENSE F0 CLEANLY TRACKED
Through the room Through the neck

The neck path never passes through the room, so the noise that destroys the air signal is simply not on it.

03 Why silence matters

It removes the behavioural barrier, not just the inconvenience.

Because the signal is captured through the neck, a murmur is enough. That is easy to mistake for a convenience feature. It is not one.

Speaking a foreign language aloud is the most anxiety-inducing act in language learning, and foreign language anxiety is among the best-documented negative predictors of achievement in the second-language acquisition literature. Adults study on commutes, in shared flats and in open-plan offices — environments in which loud, repetitive pronunciation drilling is socially impossible.

The result is an inverted allocation of effort: the skill that most requires deliberate practice is the one learners practise least. A trainer that works at murmur volume gives that practice back.

04 What makes it defensible

Six things no competitor currently does.

USP 01

Corrective haptics, not perceptual haptics

Five decades of sensory-substitution work converts incoming sound into touch so a user can perceive it. Undertone reverses the direction: it takes your outgoing production, computes the signed error, and renders the corrective action. This is the central inventive step.

USP 02

The phone as laryngeal sensor

A contact-mode signal chain matched to the tissue-conduction transfer function, with automatic contact-state detection, contact-specific gain adaptation, and opportunistic fusion with the air channel where the signal-to-noise ratio permits.

USP 03

Practice at murmur volume

Silent, private, anywhere. It removes the single largest behavioural barrier to deliberate pronunciation practice, rather than adding a convenience on top of the same loop.

USP 04

Adaptive tolerance bands

Measured accuracy re-weights the feedback bands and selects the next drill item, so the haptic vocabulary tightens as competence rises. This is what turns a feedback device into a curriculum.

USP 05

Zero hardware

Every previous attempt to bring haptics into speech training required purpose-built hardware, and that requirement is why none reached consumers at scale. Nothing to buy, ship, charge, lose or fund.

USP 06

Pronunciation as a measurement

Ordinary practice yields an objective longitudinal record: accuracy over time, most frequent error class, practice consistency. Language teaching currently has no widely deployed instrument that does this.

05 Intellectual property

Patent pending in the United Kingdom.

Two applications have been filed with the UK Intellectual Property Office, covering the two independent inventive steps at the core of the platform. A third is drafted and will be filed within the priority year.

Filed · UK IPO

Corrective Vibrotactile Feedback Method for Speech Training by Haptic Encoding of Pitch-Contour Error

Compares a learner's produced fundamental-frequency contour against a reference and renders the signed difference as a parametrically mapped vibration, so that direction, intensity envelope and temporal profile encode the corrective action required rather than a measure of correctness.

Application 1 of 3
Filed · UK IPO

Contact-Mode Extraction of Laryngeal Fundamental Frequency and Voicing Using an Unmodified Mobile Device Microphone

Captures tissue-conducted phonation with an unmodified consumer smartphone in neck contact, applying a contact-matched signal chain with contact-state detection, gain adaptation and opportunistic fusion with the air-conducted channel.

Application 2 of 3
Drafted · to file within priority year

Adaptive Haptic Tolerance-Band Control in a Pronunciation Training System

Covers the closed loop in which measured tone-production accuracy re-weights the feedback tolerance bands and selects the learner's subsequent drill items, so the haptic vocabulary progressively tightens as competence increases.

Application 3 of 3

A novelty search was commissioned before filing, not after.

A professional novelty search report was completed prior to filing, covering classifications G09B19/04, G09B21/00, G10L25/90, G10L21/06 and G06F3/01 across Espacenet, WIPO PATENTSCOPE, Google Patents and the UK IPO register, together with a non-patent-literature sweep of the sensory-substitution, vibrotactile-feedback and computer-assisted-pronunciation-training corpora.

The report confirms that the claimed combination — laryngeal pitch extraction from an unmodified handset in neck contact, coupled to corrective haptic difference feedback for second-language tone production — is not disclosed in any granted patent or published application in any jurisdiction. The closest art falls into three families, each distinguishable on the record: tactile aids for hearing-impaired users are perceptual rather than corrective; throat-microphone communications hardware uses a dedicated contact transducer rather than an unmodified handset; and computer-assisted pronunciation training returns scores through screen and audio.

Stated precisely: filed applications establish a priority date and a dated technical disclosure. They do not establish granted rights. Grant in the UK typically follows two to five years after filing, with publication at approximately eighteen months. A novelty search is not a freedom-to-operate opinion, and this page does not present it as one. The commercial moat is designed as overlapping layers — the patent family, the proprietary paired corpus, the institutional contract base — rather than as any single instrument.

06 Validation to date

What has been proven — and what has not.

Each item below states its evidential weight honestly. Confirmed commercial commitment is not blurred into general support, and internal results are labelled as internal.

Contact capture holds where air capture fails. Voiced-frame agreement with laboratory reference · internal test, small speaker panel
050100% ≈72 ≈95≈95>90 Quiet room Café level Transport level
Contact mode Air microphone, transport level

Values are approximate and are stated as the internal protocol reported them. Not peer-reviewed; the panel was small and independent replication has not yet occurred.

  • Two UK patent applications filed with the IPO, plus a completed pre-filing novelty search report, retained.
  • 142 screened UK residents surveyed across a five-week fieldwork window, recruited through six distinct community channel families — Reddit, Chinese-Forums, Facebook learner groups, Meetup, Discord and UK institutional networks.
  • 15 depth interviews with UK adult learners: nine Mandarin, two Cantonese or Vietnamese, four European-language learners at accent plateau.
  • 187 UK waitlist signups from ~1,380 sessions — 13.5% visit-to-signup, against a 6.6% all-industry and 8.4% education median, achieved with no paid acquisition.
  • Five signed Expressions of Intent from UK organisations covering all three institutional segments. Classified as qualified pipeline, not as revenue.
  • A functional MVP on TestFlight: Throat-Sense capture, four haptic error vocabularies across a forty-syllable Mandarin set, the ninety-second loop, the progress view and the teacher dashboard.
Test 01 — Technical validation

Contact capture holds where air capture fails

Pitch contours from phone-against-neck audio were compared against laboratory reference tracking (Praat autocorrelation on studio air-microphone recordings) across the full Mandarin tone inventory in three acoustic environments: quiet room, café-level noise and public-transport-level noise.

Voiced-frame agreement with the reference sat in the mid-nineties in the quiet condition, remained in the mid-nineties at café level, and stayed above 90% at transport level. The air-microphone comparator degraded sharply over the same range, falling into the low seventies as the tracker began mis-tracking on background speech.

Test 02 — Learning validation

Fourteen days, nineteen learners, two independent measures

A fourteen-day cohort study with nineteen UK-based beginner Mandarin learners recruited from the waitlist. Tone-production accuracy was measured pre and post on a held-out syllable set, scored by two native-speaker raters blinded to condition and timepoint, and independently by acoustic contour distance from native targets.

Blinded rating showed substantial improvement across the cohort, with the largest gains on the second–third tone contrast — the error class the haptic vocabularies were designed around. Inter-rater agreement was substantial on the conventional interpretation of kappa.

Limitations, stated rather than buried: nineteen participants, fourteen days, no control group, self-selected from a waitlist, run by the company itself. These are internal engineering and learning results, not peer-reviewed, and independent replication has not yet occurred. A controlled efficacy study with a UK university phonetics or applied linguistics department is a funded deliverable, budgeted at £2,500 in Year 1, £3,500 in Year 2 and £5,000 in Year 3.


07 For schools, universities and employers

The first objective instrument for pronunciation.

Language teaching currently assesses pronunciation impressionistically. Teachers judge by ear under time pressure; examination boards do the same; institutions hold no longitudinal record of any individual student's prosodic development.

Undertone's dashboard produces one automatically, as a by-product of ordinary practice: per-student, per-contrast tone accuracy over time, most frequent error class, and practice consistency. It needs no equipment beyond the phones students already carry — which means deployment is a licence, not a procurement of hardware.

More than 60% of state secondary schools report modern foreign languages recruitment difficulty. A stretched department is precisely the one most receptive to a tool that automates a measurement currently done by ear.

Year 9 Mandarin · tone accuracy Last 14 days
A. Okafor
dip-failure
78%
P. Sharma
rising-deficit
64%
L. Chen
flat-contour
91%
J. Whitfield
rising-deficit
52%
M. Rahman
falling-deficit
83%

Illustrative view. Measured tone-production accuracy and most frequent error class, generated automatically from ordinary practice — no equipment beyond the phones pupils already carry.

Institutional licenceBenchmark against the UK marketPrice
Schools per-pupil licenceFull practice loop, class grouping, per-contrast reporting UK secondary single-subject per-pupil licences observed at roughly £2–£11 per pupil per year £12 / pupil / yr
Teacher & institution dashboardSite licence, unlimited staff seats, longitudinal export Per-school unlimited-user models observed at roughly £550–£1,200 per subject per year £950 / school / yr
Corporate accent-intelligibility cohortStructured twelve-week programme, measured outcomes UK specialist studio accent programmes published at £1,350–£5,995 for five to fifteen sessions £249 / employee

Institutional sales run pilot-first, on the procurement reality the sector actually operates under: Department for Education research finds 56% of technology investment decisions are made at school level, through a process that includes an explicit piloting and trialling stage.

08 Pricing

Three buyer groups. One technical core.

The same engineering that produces contact-mode pitch extraction serves the consumer subscription, the schools licence, the corporate programme and the SDK simultaneously. Because there is no hardware, margin is pure software margin.

Learners

Free

A full daily practice loop and a limited sound set, supported by sponsored practice-break placements. A free user is inventory and word-of-mouth, not a failed conversion.

  • Ninety-second daily loopIncluded
  • Throat-Sense captureIncluded
  • Limited sound setIncluded
  • Sponsored placements£24 / '000
Join the waitlist
Learners · Premium

£9.99 / month

Ad-free, full contrast inventory, adaptive tolerance bands and the complete progress record. Priced inside the £9–£17 UK corridor for a serious pronunciation app.

  • Everything in Free, ad-free
  • Full tone & contrast inventoryIncluded
  • Language & contrast packs£14.99
  • HSK exam sprint module£29.99
Join the waitlist
Institutions & technology

From £950

Schools, universities and employers licence the measurement. Speech-technology firms and phonetics departments licence the engine and the corpus.

  • Schools per-pupil licence£12 / pupil / yr
  • Dashboard site licence£950 / yr
  • Corporate cohort£249 / employee
  • Haptic engine SDK£1,500 / month
  • Anonymised corpus licence£7,500 / yr
Talk to us

Nine streams across three buyer groups. Gross margin modelled at 91.8% in Year 1, 90.4% in Year 2 and 91.0% in Year 3 after platform commission and merchant fees. Prices shown are UK pricing and are indicative pending launch.

09 Trajectory

Each phase funded by the one before it.

Year 1 is funded by £65,000 of committed founder capital and by revenue commencing in Month 7. Years 2 to 5 are funded from reinvested operating profit, with the founder taking no salary for the first three years.

Now — Months 1–6

Build, in-house

A £24,000 build budget across eight overlapping phases: haptic pattern design and perceptual validation, native-speaker reference recording and phonetic annotation for the forty-syllable Mandarin set, security review and penetration testing, and the teacher dashboard front end. Everything that constitutes protected IP is built in-house.

Year 1 — from Month 7

Launch, retention, first pilots

Revenue commences. Focus moves to retention mechanics, funnel conversion, the first institutional pilots, and an HSK exam sprint module timed to examination sittings. R&D runs at 28% of operating expenditure.

Year 2

Tone inventories, music module, SDK

Cantonese, Vietnamese and Thai tone inventories. The music and rhythm module at full scope. SDK packaging that makes the C++ core licensable to third parties. The controlled efficacy study with a university partner. R&D rises to 33%.

Year 3

Non-tonal contrast packs, Ireland

Four precision contrast packs — Spanish trilled R, French nasal vowels, Arabic emphatics, Japanese vowel length. International expansion begins with Ireland. R&D at 38%, headcount at eleven.

Years 4–5

Germany, the Netherlands, the US and Australia

Continued expansion, with the intended outcome a strategic acquisition in the five-to-seven-year window — valued on the patent family, the proprietary paired corpus and the institutional contract base rather than on consumer subscriber count alone.


LanguagesFounder
RussianNative
TatarNative
EnglishFluent
MandarinElementary

Four years teaching English to adults in London. The fourth line on this list is where the problem was found.

10 The founder

I could hear that my tones were wrong when a native speaker told me so. I could not hear how they were wrong — and every application answered a failed attempt by playing the correct version again, as though the problem were that I had not listened carefully enough.

Daliia Gaskarova speaks Russian and Tatar natively, English fluently, and Chinese at an elementary level. She spent four years teaching English to adults in London, watching the same wall from the other side: Russian speakers who could not produce the English th, Arabic speakers who could not separate p from b. Learning English is the reason she was able to move to London, study, and build a career here — and she is acutely aware that the people around her who did not acquire it are living different lives for no other reason.

Daliia Gaskarova Founder & Director · London, United Kingdom

11 Early access

Join 187 UK learners on the waitlist.

The MVP is on TestFlight and beta places are allocated from the waitlist. Tell us the language you're working on and we'll match you to the right cohort.

Raw voice audio is processed on your own device. Undertone is an educational training platform, not a clinical instrument: it measures pronunciation accuracy for the purpose of language learning, and does not diagnose, screen, treat or monitor speech, hearing or developmental conditions.