Most people understand English better than they speak it. When it comes to actually talking, fear of judgment, lack of practice, or just awkward silence gets in the way.
Vocao is an AI tutor built for speaking practice. It gives you a safe, low-pressure space to talk — through simple, playful conversations
I could understand English just fine. Until I had to speak — especially during interviews. My mind went blank. Words vanished. Confidence politely left the room. At some point I thought: “Why is it so hard to start speaking… nicely?” That question turned into Vocao


How might we help English learners practice speaking naturally — without fear or judgment? The challenge was not just to design a learning tool — but to design an emotionally safe, motivating experience that makes people want to talk
I was the co-foinder and solo designer. I owned:
• User research
• UX & product flows
• Visual direction
• Illustrations & animations
• Iterative design process
We began small — talking to friends, students, and eventually our real target users: international learners, TOEFL/IELTS takers, professionals building global careers, and curious language enthusiasts.
Through user interviews and competitive analysis (including a chat with a former Duolingo engineer), we uncovered several patterns:
Many felt anxious speaking with real tutors
Many found lessons too expensive or inconsistent
Most lacked fun, daily motivation
We also found that some learners were hesitant to speak with AI. So our goal became clear — make the experience so interactive, playful, and human-like that we might gently convert at least some of these skeptics


To ship an MVP, we applied a strict filter: Does this feature directly help the user speak more?
Based on that principle, the MVP included four essential components:
Topic-based conversations
Free practice mode
Daily streak tracking
Post-conversation feedback
Everything else — avatars, advanced leveling, personalization, challenges, topic packs — was deliberately postponed. This allowed us to build fast, test real behavior, and avoid premature complexity
Designing a voice-first product meant designing for silence, hesitation, and fear.
Key decisions:
Clear session states (listening / thinking / waiting — no confusion)
Gentle prompts instead of error messages
Friendly endings that reward effort, not perfection
The UI never asks: “Why did you say that?”
It says: “Nice try. Let’s keep going.”


To make the product emotionally engaging, I explored a visual language that felt warm, nostalgic, and futuristic. Many users grew up with pixel-style games and associated them with comfort and curiosity. Pure pixel UI felt too flat for an AI product, so I created a hybrid “retro-futuristic” design:
AI-generated backgrounds evoked soft, dreamlike environments
Hand-drawn pixel characters added familiarity and charm
Everything was designed using Figma, ChatGPT, and Midjourney
After the designs were finalized, I built a prototype to see how the interface would feel in motion.
We tested it with a few users, gathered initial feedback, made improvements, and then launched an early version
From week 2 onward, we shipped new builds almost every week, iterating based on:
Conversation issues
UX issues
Technical issues
