Skip to main content

Separation of Word Clusters / Interactive Morphology Tab

Often times for several Asian languages, two or more words will automatically be grouped into a single block (without manual merging). However, it is not possible to break up these clusters, for example to look at a single word isolated or to save it without the other word attached.

The morphology tab already shows the two words separately, but then does not allow for interaction with the words to save them individually.

This is even more problematic with the use of grammatical particles behind the word. A fix or addition through a new feature would be life changing!!

Status: Completed14 comments

Log in to comment and vote

Comments14

  • mrtn-f

    •

    Dec 22, 2025

    This can even happen between different sentences and different speakers, where the last words of the prior sentence and the first word of the next sentence are automatically clustered.

    In the image below, the “What” is already part of the next sentence and by a different speaker, but the entire thing is recognized as an atomic word.

  • Sasha

    Team•

    Dec 22, 2025

    This is indeed very weird.

    Would you be able to show me a sentence where the tokenization is wrong and what the tokenization should be?

    The tokenization is the act of separating characters in a sentence into single units.

    • mrtn-f

      •

      Dec 22, 2025

      Here, for example, almost the entire sentence is merged into a single token.

      A correct tokenization would be:
      (여기) (촬영 하는) (구나) (영화) (예요)

      In the second bracket there is a verb that means “filming” but it translates roughly to “making/doing film” so there one could also separate the token once more but it is technically just a compound verb.

      The Morphology tab already mostly identifies the correct tokens (with verbs in their dictionary form). So a temporary fix would also be to enable the interaction with the components in the Morphology tab. There, one could then use the usual functionality like getting examples for usage or saving them to the notebook.

      • mrtn-f

        •

        Dec 22, 2025

        Or alternatively, similar to the way we merge tokens, the ability to select/disconnect them even for tokens that are seen as atomic would be amazing.

        • Sasha

          Team•

          Dec 22, 2025

          gotcha. Im pushing an update soon that should adress this issue

          • mrtn-f

            •

            Dec 23, 2025

            Insane!!! So positively surprised how quickly that was addressed. Amazing work. Instantly won me over 🤝

            • Sasha

              Team•

              Dec 23, 2025

              Does it work better now ? Do you have an example ?

              • mrtn-f

                •

                Dec 23, 2025

                Haven’t checked yet, but will do today or tomorrow. But I was commenting on the speed at which issues and feedback are addressed here. Truly amazing!

                • Sasha

                  Team•

                  Dec 23, 2025

                  Great to hear 👊🏼

                  • mrtn-f

                    •

                    Dec 27, 2025

                    After the update, I still experience the same issue.

                    I obviously don’t know the code base, but maybe making the morphology tab interactive so that one can click on the words and save them / read up on their use cases etc. would be easier. At least, it would be massively helpful for not only this issue but language learning in general.

                    • Sasha

                      Team•

                      Dec 27, 2025

                      I’m getting for instance

                      On firefox version 2025.194

                      are you testing on that version ?

                      • mrtn-f

                        •

                        Dec 27, 2025

                        Yes i am testing on 2025.194.

                        It doesn’t happen on every sentence but around 5-10% of the time.

                        It might also be different for non-korean shows, but I am not sure.

                        At least, I only tested it for native korean shows.

                        • Sasha

                          Team•

                          Dec 27, 2025

                          So after investigation, I was able to confirm your issue, and those are the solutions that claude opus 4.5 is suggesting.


                          Solution A: AI/LLM-Based Tokenization (Backend)

                          Description: Use the existing AI inference infrastructure to tokenize Korean text server-side. Extend the SENTENCE_MATCHING_PROMPT_TEXT with Korean-specific rules.

                          Pros:

                          • High accuracy for morphological analysis

                          • Already have infrastructure (RevampedInferenceService)

                          • Handles archaic/complex Korean well

                          • No additional dependencies

                          • Consistent with existing Japanese tokenization approach

                          Cons:

                          • Requires network request (latency ~200-500ms)

                          • API cost per tokenization

                          • Requires authentication

                          • Not available offline

                          O Complexity: O(n) where n = text length (single API call)


                          Solution B: WebAssembly Korean Morphological Analyzer (khaiii.js)

                          Description: Use https://github.com/puilp0502/khaiii.js, a WASM port of Korea's Kakao Hangul Input Interface, for client-side morphological analysis.

                          Pros:

                          • High accuracy (trained on Korean corpus)

                          • Client-side (no latency after initial load)

                          • No API costs

                          • Works offline

                          Cons:

                          • Large WASM bundle (~10-20MB dictionary)

                          • Initial load time

                          • Maintenance risk (limited community support)

                          • May not handle archaic Korean well

                          O Complexity: O(n) with dictionary lookup overhead


                          Solution C: Heuristic Korean Syllable-Based Segmentation

                          Description: Implement rule-based segmentation using Korean syllable structure patterns. Split on common particle boundaries (은/는/이/가/을/를/에/로/의/와/과) and verb endings.

                          Pros:

                          • No dependencies

                          • Fast (pure JS)

                          • Works offline

                          • Small bundle size

                          • No API costs

                          Cons:

                          • Lower accuracy than ML-based solutions

                          • Doesn't understand context

                          • Can over-segment or under-segment

                          • Requires maintaining rule database

                          O Complexity: O(n) with regex matching


                          Solution D: Hybrid Approach (Heuristic + AI Fallback)

                          Description: Use heuristic segmentation client-side for immediate display, then optionally refine with AI-based tokenization when available/needed.

                          Pros:

                          • Best of both worlds

                          • Fast initial display

                          • High accuracy when refined

                          • Graceful degradation

                          • Works offline with fallback

                          Cons:

                          • More complex implementation

                          • Potential UI "jump" when tokens refine

                          • Two code paths to maintain

                          O Complexity: O(n) + optional O(n) API call


                          Solution E: Backend-First Tokenization (Pre-tokenized Subtitles)

                          Description: Modify the subtitle pipeline to always tokenize on backend before sending to frontend. Korean subtitles arrive pre-tokenized.

                          Pros:

                          • Cleanest separation of concerns

                          • High accuracy

                          • Frontend stays simple

                          • Consistent with existing neural transcription flow

                          Cons:

                          • Requires backend changes

                          • All Korean content needs re-processing

                          • Higher backend load

                          • Breaks existing cached content

                          O Complexity: O(n) per subtitle segment


                          Solution F: Extended Intl.Segmenter with Post-Processing

                          Description: Keep Intl.Segmenter but add post-processing to split oversized Korean tokens using particle detection patterns.

                          Pros:

                          • Minimal changes to existing code

                          • Uses native browser API

                          • Fast

                          • No dependencies

                          Cons:

                          • Limited accuracy

                          • Pattern maintenance

                          • Won't handle all cases

                          O Complexity: O(n)


                          I guess I will go with LLM tokenisation like we do for Japanese and see if it helps.

                          • mrtn-f

                            •

                            Dec 28, 2025

                            Sounds good.

                            At least C) and D) sound like a load of work and might still be less accurate because of custom heuristics that need to be found.

                            E) likely requires a lot of reworking, depending on how your backend looks (probably not worth it or realistic). I don’t really get approach F) actually, but both A and B seem good.

                            I looked at the github repo for B) and it seems pretty straight forward. Depending on how much you can buffer, preloading the tokenizer tool on start up when using Korean should make it very fast when doing live tokenization (likely faster than the LLM approach, but can’t tell without testing). But if you already apply a similar method to A) for Japanese, i guess that is definitely worth a try.

                            I am also pretty sure, that LLMs will perfectly tokenize Korean sentences. At most, maybe with a bit of an inconsistency as to whether particles belong to the word before or are seen as separate (and similar cases). But either you establish a context for the LLM that defines those cases and sets concrete rules (could also let the LLM decide on that ones and use it as context), or you just let it be. Even if there is a bit of inconsistency, that shouldn’t really matter that much