Unicode character ſ is matched as itself and as 's.' #14294

diffsetter · 2024-03-25T11:13:48Z

Steps to reproduce

Open a utf8 encoded file in vim containing the line "Die Gleichheit fordert das Nachdenken heraus durch Fragen, die ſich daran knüpfen und nicht ganz leicht zu beantworten ſind."
use :set ignorecase
replace the character ſ by s using :%s/\%u017F/s/g

Expected behaviour

The two occurrences of 'ſ' will be replaced by 's' resulting in "Die Gleichheit fordert das Nachdenken heraus durch Fragen, die sich daran knüpfen und nicht ganz leicht zu beantworten sind." However, the actual result is "Die Gleichheit fordert dasNachdenken herausdurch Fragen, die sich daran knüpfen und nicht ganz leicht zu beantworten sind.", i.e., the two original s characters together with the following character are also replaced by 's' as if I had used the command :%s/s./s/g. See also this discussion.

Version of Vim

9.1.151 but also older like 8.0

Environment

system: x86_64 GNU/Linux
terminal: konsole, linux
$TERM: linux, xterm-256color
$LANG: de_DE.UTF-8, en_GB.UTF-8, C.UTF-8

Logs and stack traces

No response

The text was updated successfully, but these errors were encountered:

chrisbra · 2024-03-25T15:12:14Z

hm, it works with :set regexpengine=1

diffsetter · 2024-03-25T16:25:56Z

hm, it works with :set regexpengine=1

Indeed, it does. But does that mean it's not a bug? I didn't know that option. The help file says: "Note that when using the NFA engine [the one that must have been chosen by vim automatically in this case] and the pattern contains something that is not supported the pattern will not match…" But here something is matched that shouldn't have been matched.

RestorerZ · 2024-03-25T16:42:07Z

It kind of reminds me of this. issues #12579

chrisbra · 2024-03-26T17:01:40Z

yes and also related: #13682

…ngle byte char When the regexp engine compares two utf-8 codepoints case insensitively it may match an adjacent character, because it assumes it can step over as many bytes as the pattern contains. Let's consider the pattern 'ſ' and the string 's '. When comparing and ignoring case, the single character 's' matches, and since it matches Vim will try to step over the match (by the amount of bytes of the pattern), assuming that since it matches, the length of both strings is the same. However in that case, it should only step over the single byte value 's' by 1 byte and try to start matching after it again. So for the backtracking engine we need to ensure: - we try to match the correct length for the pattern and the text - in case of a match, we step over it correctly For the NFA engine, it basically already handles it already. Here we just need to make sure, that find_match_text is only called if prog->match_text is actually not only NULL but also doesn't point to 0x0 (NUL). fixes: vim#14294 Signed-off-by: Christian Brabandt <cb@256bit.org>

chrisbra · 2024-04-06T19:26:54Z

I checked it and I think I found the issue. However, I have a question. Can I assume correctly, that 'ſ' should match the lower case 's' if ignoring case?

…te char When the regexp engine compares two utf-8 codepoints case insensitively it may match an adjacent character, because it assumes it can step over as many bytes as the pattern contains. This however is not necessarily true because of case-folding, a multi-byte UTF-8 character can be considered equal to some single-byte value. Let's consider the pattern 'ſ' and the string 's'. When comparing and ignoring case, the single character 's' matches, and since it matches Vim will try to step over the match (by the amount of bytes of the pattern), assuming that since it matches, the length of both strings is the same. However in that case, it should only step over the single byte value 's' by 1 byte and try to start matching after it again. So for the backtracking engine we need to ensure: - we try to match the correct length for the pattern and the text - in case of a match, we step over it correctly The same thing can happen for the NFA engine, when skipping to the next character to test for a match. We are skipping over the regstart pointer, however we do not consider the case that because of case-folding we may need to adjust the number of bytes to skip over. So this needs to be adjusted in find_match_text() as well. A related issue turned out, when prog->match_text is actually empty. In that case we should try to find the next match and skip this condition. fixes: vim#14294 Signed-off-by: Christian Brabandt <cb@256bit.org>

Problem: Regex engines do not handle case-folding well Solution: Correctly calculate byte length of characters to skip When the regexp engine compares two utf-8 codepoints case insensitively it may match an adjacent character, because it assumes it can step over as many bytes as the pattern contains. This however is not necessarily true because of case-folding, a multi-byte UTF-8 character can be considered equal to some single-byte value. Let's consider the pattern 'ſ' and the string 's'. When comparing and ignoring case, the single character 's' matches, and since it matches Vim will try to step over the match (by the amount of bytes of the pattern), assuming that since it matches, the length of both strings is the same. However in that case, it should only step over the single byte value 's' so by 1 byte and try to start matching after it again. So for the backtracking engine we need to ensure: - we try to match the correct length for the pattern and the text - in case of a match, we step over it correctly The same thing can happen for the NFA engine, when skipping to the next character to test for a match. We are skipping over the regstart pointer, however we do not consider the case that because of case-folding we may need to adjust the number of bytes to skip over. So this needs to be adjusted in find_match_text() as well. A related issue turned out, when prog->match_text is actually empty. In that case we should try to find the next match and skip this condition. fixes: vim/vim#14294 closes: vim/vim#14433 vim/vim@7a27c10 Co-authored-by: Christian Brabandt <cb@256bit.org>

…28259) Problem: Regex engines do not handle case-folding well Solution: Correctly calculate byte length of characters to skip When the regexp engine compares two utf-8 codepoints case insensitively it may match an adjacent character, because it assumes it can step over as many bytes as the pattern contains. This however is not necessarily true because of case-folding, a multi-byte UTF-8 character can be considered equal to some single-byte value. Let's consider the pattern 'ſ' and the string 's'. When comparing and ignoring case, the single character 's' matches, and since it matches Vim will try to step over the match (by the amount of bytes of the pattern), assuming that since it matches, the length of both strings is the same. However in that case, it should only step over the single byte value 's' so by 1 byte and try to start matching after it again. So for the backtracking engine we need to ensure: - we try to match the correct length for the pattern and the text - in case of a match, we step over it correctly The same thing can happen for the NFA engine, when skipping to the next character to test for a match. We are skipping over the regstart pointer, however we do not consider the case that because of case-folding we may need to adjust the number of bytes to skip over. So this needs to be adjusted in find_match_text() as well. A related issue turned out, when prog->match_text is actually empty. In that case we should try to find the next match and skip this condition. fixes: vim/vim#14294 closes: vim/vim#14433 vim/vim@7a27c10 Co-authored-by: Christian Brabandt <cb@256bit.org>

…te char When the regexp engine compares two utf-8 codepoints case insensitively it may match an adjacent character, because it assumes it can step over as many bytes as the pattern contains. This however is not necessarily true because of case-folding, a multi-byte UTF-8 character can be considered equal to some single-byte value. Let's consider the pattern 'ſ' and the string 's'. When comparing and ignoring case, the single character 's' matches, and since it matches Vim will try to step over the match (by the amount of bytes of the pattern), assuming that since it matches, the length of both strings is the same. However in that case, it should only step over the single byte value 's' by 1 byte and try to start matching after it again. So for the backtracking engine we need to ensure: - we try to match the correct length for the pattern and the text - in case of a match, we step over it correctly The same thing can happen for the NFA engine, when skipping to the next character to test for a match. We are skipping over the regstart pointer, however we do not consider the case that because of case-folding we may need to adjust the number of bytes to skip over. So this needs to be adjusted in find_match_text() as well. A related issue turned out, when prog->match_text is actually empty. In that case we should try to find the next match and skip this condition. Note: this breaks Mail and CSS Syntax highlighting and CI on FreeBSD/MacOS vim#14487 and https://groups.google.com/d/msgid/vim_dev/CAJkCKXtui%3DDTWx9eV8Dbs19XoFL9b63ObSNXWCRvLsEZCB6yfw%40mail.gmail.com. fixes: vim#14294 Signed-off-by: Christian Brabandt <cb@256bit.org>

…te char This is v2 of v9.1.296, but still draft to see if this still breaks. When the regexp engine compares two utf-8 codepoints case insensitively it may match an adjacent character, because it assumes it can step over as many bytes as the pattern contains. This however is not necessarily true because of case-folding, a multi-byte UTF-8 character can be considered equal to some single-byte value. Let's consider the pattern 'ſ' and the string 's'. When comparing and ignoring case, the single character 's' matches, and since it matches Vim will try to step over the match (by the amount of bytes of the pattern), assuming that since it matches, the length of both strings is the same. However in that case, it should only step over the single byte value 's' by 1 byte and try to start matching after it again. So for the backtracking engine we need to ensure: - we try to match the correct length for the pattern and the text - in case of a match, we step over it correctly There is one tricky thing for the backtracing engine. We also need to calculate correctly the number of bytes to compare the 2 different utf-8 strings s1 and s2. So we will count the number of characters in s1 that the byte len specified. Then we count the number of bytes to step over the same number of characters in string s2 and then we can correctly compare the 2 utf-8 strings. A similar thing can happen for the NFA engine, when skipping to the next character to test for a match. We are skipping over the regstart pointer, however we do not consider the case that because of case-folding we may need to adjust the number of bytes to skip over. So this needs to be adjusted in find_match_text() as well. A related issue turned out, when prog->match_text is actually empty. In that case we should try to find the next match and skip this condition. Note: this breaks Mail and CSS Syntax highlighting and CI on FreeBSD/MacOS vim#14487 and https://groups.google.com/d/msgid/vim_dev/CAJkCKXtui%3DDTWx9eV8Dbs19XoFL9b63ObSNXWCRvLsEZCB6yfw%40mail.gmail.com. fixes: vim#14294 Signed-off-by: Christian Brabandt <cb@256bit.org>

diffsetter added the bug label Mar 25, 2024

chrisbra mentioned this issue Apr 6, 2024

regexp: wrong match when 'ic' is set and comparing multi-byte with single byte char #14433

Closed

chrisbra closed this as completed in 7a27c10 Apr 9, 2024

zeertzjq mentioned this issue Apr 9, 2024

vim-patch:9.1.0296: regexp: engines do not handle case-folding well neovim/neovim#28259

Merged

chrisbra reopened this Apr 10, 2024

chrisbra linked a pull request May 12, 2024 that will close this issue

regexp: wrong match with 'ic' and comparing multi-byte with single byte char #14756

Draft

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Unicode character ſ is matched as itself and as 's.' #14294

Unicode character ſ is matched as itself and as 's.' #14294

diffsetter commented Mar 25, 2024

chrisbra commented Mar 25, 2024

diffsetter commented Mar 25, 2024

RestorerZ commented Mar 25, 2024

chrisbra commented Mar 26, 2024

chrisbra commented Apr 6, 2024 •

edited

Unicode character ſ is matched as itself and as 's.' #14294

Unicode character ſ is matched as itself and as 's.' #14294

Comments

diffsetter commented Mar 25, 2024

Steps to reproduce

Expected behaviour

Version of Vim

Environment

Logs and stack traces

chrisbra commented Mar 25, 2024

diffsetter commented Mar 25, 2024

RestorerZ commented Mar 25, 2024

chrisbra commented Mar 26, 2024

chrisbra commented Apr 6, 2024 • edited

chrisbra commented Apr 6, 2024 •

edited