I'm building a Chinese learning project called HanziHero, and one feature started with what looked like a ridiculously simple data problem:

Why do learners keep mixing up Chinese characters they already know?

Think:

未 / 末

土 / 士