Post ·
Face, Voice and a Pass-Phrase: Lessons From a 2018 Hackathon
One short video, three checks: face, voice and the spoken pass-phrase. What a first-place 2018 Microsoft/TAG hackathon taught me about layered authentication.


On October 13, 2018, my team took first place at a Microsoft and TAG hackathon in Atlanta. Our challenge was a multimodal biometrics system: confirm who someone is from one short video, by matching their face, their voice and the pass-phrase they speak.
The brief
Every team member enrolled first, recording a video while saying one pass-phrase from a list. Verification took a recording, a user ID and the pass-phrase, and returned pass or fail. The judges scored four things: enrolling everyone, verifying everyone, resisting fraudulent attempts, and the polish of the experience. Teams earned extra points for demonstrating attacks and showing the system turn them away.
The pass-phrase list looked whimsical (“apple juice tastes funny after toothpaste”), but it wasn't arbitrary. It was the ten English phrases that Microsoft's text-dependent speaker verification supported.
One recording, three checks
The useful move was to stop treating the video as one input. It carries three independent signals, so split it early and judge each one on its own:
- Frames pulled from the video, compared against the enrolled face.
- The audio track, extracted to a plain WAV file, compared against the enrolled voice.
- The words themselves, which have to match the pass-phrase the user claims.
Then decide once, at the end: all three pass, or the attempt fails. A face alone can be a printed photo. A voice alone can be a recording. A pass-phrase alone is a password said out loud, which is roughly the opposite of a secret.
What carried over
- Each factor is weak alone. The strength comes from requiring them together and keeping them independent, so one spoofed signal can't vouch for the others.
- Test the attack, not just the login. The judges asked teams to show fraudulent attempts failing, which is the right instinct for any authentication design. A demo that only shows the happy path proves the happy path.
- Enrollment is the product. Verification can only be as good as the samples captured on day one, and poor lighting at enrollment is a bug that ships forever.
- Put each check behind its own seam. The services behind the signals changed far more than the pattern did.
Looking back from 2026
The pattern aged better than the parts. The starter notebook asked the face service for age, gender and emotion along with the face itself. In June 2022 Microsoft retired those attributes and put facial identification and verification behind an application process. Speaker Recognition has since been retired altogether.
Cloned voices and synthetic video have also made “looks and sounds like you” a weaker signal than it was in 2018. I'd still layer independent checks, but today at least one of them would be something you have, like a device or a passkey, with face and voice as supporting evidence rather than the gate.
My prize was an Xbox One. The next day my son won a LEGO Technic Bugatti at a LEGO show at the Cobb Galleria. His prize was slightly bigger than mine.
A month later I was back at another hackathon, this time building an ATM out of a Raspberry Pi. That story is in A Raspberry Pi ATM: Lessons From a 2018 FinTech Hackathon.