aboutsummaryrefslogtreecommitdiff
path: root/src/include/regex/regex.h
diff options
context:
space:
mode:
authorThomas Munro <tmunro@postgresql.org>2024-07-05 15:25:31 +1200
committerThomas Munro <tmunro@postgresql.org>2024-07-05 15:25:31 +1200
commit18501841bcb4e693b9f1e9da2b2fb524c78940d8 (patch)
tree77e63b4f9b9b857a72bee82c88eefb755052ac1f /src/include/regex/regex.h
parent1eff8279d494b96f0073df78abc74954a2f6ee54 (diff)
downloadpostgresql-18501841bcb4e693b9f1e9da2b2fb524c78940d8.tar.gz
postgresql-18501841bcb4e693b9f1e9da2b2fb524c78940d8.zip
Add simple codepoint redirections to unaccent.rules.
Previously we searched for code points where the Unicode data file listed an equivalent combining character sequence that added accents. Some codepoints redirect to a single other codepoint, instead of doing any combining. We can follow those references recursively to get the answer. Per bug report #18362, which reported missing Ancient Greek characters. Specifically, precomposed characters with oxia (from the polytonic accent system used for old Greek) just point to precomposed characters with tonos (from the monotonic accent system for modern Greek), and we have to follow the extra hop to find out that they are composed with an acute accent. Besides those, the new rule also: * pulls in a lot of 'Mathematical Alphanumeric Symbols', which are copies of the Latin and Greek alphabets and numbers rendered in different typefaces, and * corrects a single mathematical letter that previously came from the CLDR transliteration file, but the new rule extracts from the main Unicode database file, where clearly the latter is right and the former is a wrong (reported to CLDR). Reported-by: Cees van Zeeland <cees.van.zeeland@freedom.nl> Reviewed-by: Robert Haas <robertmhaas@gmail.com> Reviewed-by: Peter Eisentraut <peter@eisentraut.org> Reviewed-by: Michael Paquier <michael@paquier.xyz> Discussion: https://postgr.es/m/18362-be6d0cfe122b6354%40postgresql.org
Diffstat (limited to 'src/include/regex/regex.h')
0 files changed, 0 insertions, 0 deletions