Separation of Word Clusters / Interactive Morphology Tab

Often times for several Asian languages, two or more words will automatically be grouped into a single block (without manual merging). However, it is not possible to break up these clusters, for example to look at a single word isolated or to save it without the other word attached.
The morphology tab already shows the two words separately, but then does not allow for interaction with the words to save them individually.
This is even more problematic with the use of grammatical particles behind the word. A fix or addition through a new feature would be life changing!!
Log in to comment and vote
Comments14
mrtn-f
Dec 22, 2025
This can even happen between different sentences and different speakers, where the last words of the prior sentence and the first word of the next sentence are automatically clustered.
In the image below, the “What” is already part of the next sentence and by a different speaker, but the entire thing is recognized as an atomic word.
Sasha
Dec 22, 2025
This is indeed very weird.
Would you be able to show me a sentence where the tokenization is wrong and what the tokenization should be?
The tokenization is the act of separating characters in a sentence into single units.
mrtn-f
Dec 22, 2025
Here, for example, almost the entire sentence is merged into a single token.
A correct tokenization would be:
(여기) (촬영 하는) (구나) (영화) (예요)
In the second bracket there is a verb that means “filming” but it translates roughly to “making/doing film” so there one could also separate the token once more but it is technically just a compound verb.
The Morphology tab already mostly identifies the correct tokens (with verbs in their dictionary form). So a temporary fix would also be to enable the interaction with the components in the Morphology tab. There, one could then use the usual functionality like getting examples for usage or saving them to the notebook.
mrtn-f
Dec 22, 2025
Or alternatively, similar to the way we merge tokens, the ability to select/disconnect them even for tokens that are seen as atomic would be amazing.
Sasha
Dec 22, 2025
gotcha. Im pushing an update soon that should adress this issue
mrtn-f
Dec 23, 2025
Insane!!! So positively surprised how quickly that was addressed. Amazing work. Instantly won me over 🤝
Sasha
Dec 23, 2025
Does it work better now ? Do you have an example ?
mrtn-f
Dec 23, 2025
Haven’t checked yet, but will do today or tomorrow. But I was commenting on the speed at which issues and feedback are addressed here. Truly amazing!
Sasha
Dec 23, 2025
Great to hear 👊🏼
mrtn-f
Dec 27, 2025
After the update, I still experience the same issue.
I obviously don’t know the code base, but maybe making the morphology tab interactive so that one can click on the words and save them / read up on their use cases etc. would be easier. At least, it would be massively helpful for not only this issue but language learning in general.
Sasha
Dec 27, 2025
I’m getting for instance
On firefox version 2025.194
are you testing on that version ?
mrtn-f
Dec 27, 2025
Yes i am testing on 2025.194.
It doesn’t happen on every sentence but around 5-10% of the time.
It might also be different for non-korean shows, but I am not sure.
At least, I only tested it for native korean shows.
Sasha
Dec 27, 2025
So after investigation, I was able to confirm your issue, and those are the solutions that claude opus 4.5 is suggesting.
Solution A: AI/LLM-Based Tokenization (Backend)
Description: Use the existing AI inference infrastructure to tokenize Korean text server-side. Extend the SENTENCE_MATCHING_PROMPT_TEXT with Korean-specific rules.
Pros:
High accuracy for morphological analysis
Already have infrastructure (RevampedInferenceService)
Handles archaic/complex Korean well
No additional dependencies
Consistent with existing Japanese tokenization approach
Cons:
Requires network request (latency ~200-500ms)
API cost per tokenization
Requires authentication
Not available offline
O Complexity: O(n) where n = text length (single API call)
Solution B: WebAssembly Korean Morphological Analyzer (khaiii.js)
Description: Use https://github.com/puilp0502/khaiii.js, a WASM port of Korea's Kakao Hangul Input Interface, for client-side morphological analysis.
Pros:
High accuracy (trained on Korean corpus)
Client-side (no latency after initial load)
No API costs
Works offline
Cons:
Large WASM bundle (~10-20MB dictionary)
Initial load time
Maintenance risk (limited community support)
May not handle archaic Korean well
O Complexity: O(n) with dictionary lookup overhead
Solution C: Heuristic Korean Syllable-Based Segmentation
Description: Implement rule-based segmentation using Korean syllable structure patterns. Split on common particle boundaries (은/는/이/가/을/를/에/로/의/와/과) and verb endings.
Pros:
No dependencies
Fast (pure JS)
Works offline
Small bundle size
No API costs
Cons:
Lower accuracy than ML-based solutions
Doesn't understand context
Can over-segment or under-segment
Requires maintaining rule database
O Complexity: O(n) with regex matching
Solution D: Hybrid Approach (Heuristic + AI Fallback)
Description: Use heuristic segmentation client-side for immediate display, then optionally refine with AI-based tokenization when available/needed.
Pros:
Best of both worlds
Fast initial display
High accuracy when refined
Graceful degradation
Works offline with fallback
Cons:
More complex implementation
Potential UI "jump" when tokens refine
Two code paths to maintain
O Complexity: O(n) + optional O(n) API call
Solution E: Backend-First Tokenization (Pre-tokenized Subtitles)
Description: Modify the subtitle pipeline to always tokenize on backend before sending to frontend. Korean subtitles arrive pre-tokenized.
Pros:
Cleanest separation of concerns
High accuracy
Frontend stays simple
Consistent with existing neural transcription flow
Cons:
Requires backend changes
All Korean content needs re-processing
Higher backend load
Breaks existing cached content
O Complexity: O(n) per subtitle segment
Solution F: Extended Intl.Segmenter with Post-Processing
Description: Keep Intl.Segmenter but add post-processing to split oversized Korean tokens using particle detection patterns.
Pros:
Minimal changes to existing code
Uses native browser API
Fast
No dependencies
Cons:
Limited accuracy
Pattern maintenance
Won't handle all cases
O Complexity: O(n)
I guess I will go with LLM tokenisation like we do for Japanese and see if it helps.
mrtn-f
Dec 28, 2025
Sounds good.
At least C) and D) sound like a load of work and might still be less accurate because of custom heuristics that need to be found.
E) likely requires a lot of reworking, depending on how your backend looks (probably not worth it or realistic). I don’t really get approach F) actually, but both A and B seem good.
I looked at the github repo for B) and it seems pretty straight forward. Depending on how much you can buffer, preloading the tokenizer tool on start up when using Korean should make it very fast when doing live tokenization (likely faster than the LLM approach, but can’t tell without testing). But if you already apply a similar method to A) for Japanese, i guess that is definitely worth a try.
I am also pretty sure, that LLMs will perfectly tokenize Korean sentences. At most, maybe with a bit of an inconsistency as to whether particles belong to the word before or are seen as separate (and similar cases). But either you establish a context for the LLM that defines those cases and sets concrete rules (could also let the LLM decide on that ones and use it as context), or you just let it be. Even if there is a bit of inconsistency, that shouldn’t really matter that much