Skip to content

Use autoCaptureLanguage to analyze user profile - #317

Open
berteauxjb wants to merge 3 commits into
tickernelz:mainfrom
berteauxjb:fix/profile-language-config
Open

berteauxjb wants to merge 3 commits into
tickernelz:mainfrom
berteauxjb:fix/profile-language-config

Conversation

@berteauxjb

Copy link
Copy Markdown

The prompt for the user profile mentions that the model should detect the language of the user. In my case, the model detected Catalan even though my prompts were written in English. After further investigation, I realized that the setting autoCaptureLanguage was only used to generate memories, but not to generate the user profile.

This PR mirrors the implementation of the language selection already implemented for the memories, but for the user profile.

Disclaimer: the issue was investigated and fixed by Claude Sonnet 5.

User profile can randomly be generated in another language than what the
user uses. There's already a mechanism to select the language of the
memories, this commit uses it in the user profile as well.
@karaaslanz

Copy link
Copy Markdown

I traced the new resolveProfileLanguageName() call against the existing auto-capture path. There is one correctness gap in the "auto" case: auto-capture runs detectLanguage(userPrompt), but this PR runs detectLanguage(context). The profile context is not just user text — buildUserAnalysisContext() wraps the prompts in a fairly large English instruction scaffold (# User Profile Analysis, analysis guidelines, profile-size/category guidance, etc.) and may also include existing profile text. That can bias detection toward English, especially for short non-English prompts, so the original wrong-language symptom can still occur even though the config is now consulted.
I would keep the configured-language path as-is, but for auto detect from the actual recent-prompt corpus before it is wrapped in the analysis context (or pass that corpus separately into analyzeUserProfile). That also mirrors the existing auto-capture semantics more closely. A focused regression where short non-English prompts are surrounded by the current English scaffold would lock this down.

Previous commit was detecting the language based on the user prompts
wrapped in the analysis context, which contains English instructions.
This could skew the language detection to English. This commit addresses
the issue by only analyzing the user prompts.
@berteauxjb

Copy link
Copy Markdown
Author

I traced the new resolveProfileLanguageName() call against the existing auto-capture path. There is one correctness gap in the "auto" case: auto-capture runs detectLanguage(userPrompt), but this PR runs detectLanguage(context). The profile context is not just user text — buildUserAnalysisContext() wraps the prompts in a fairly large English instruction scaffold (# User Profile Analysis, analysis guidelines, profile-size/category guidance, etc.) and may also include existing profile text. That can bias detection toward English, especially for short non-English prompts, so the original wrong-language symptom can still occur even though the config is now consulted. I would keep the configured-language path as-is, but for auto detect from the actual recent-prompt corpus before it is wrapped in the analysis context (or pass that corpus separately into analyzeUserProfile). That also mirrors the existing auto-capture semantics more closely. A focused regression where short non-English prompts are surrounded by the current English scaffold would lock this down.

Thanks for the feedback! I updated the PR to:

  • only run the language analysis on the user prompts (without system context)
  • add a unit test to exercise the changes

@karaaslanz karaaslanz left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-checked the updated head after my earlier feedback. The language detection now runs against the actual user-prompt corpus rather than the English analysis scaffold, and the added regression coverage exercises that exact failure mode.

That addresses the correctness concern from my previous review scope. Looks good from this scope.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants