Skip to content

Compact the _Unicode_property_data tables. - #2757

Merged
Stephan T. Lavavej (StephanTLavavej) merged 2 commits into
microsoft:mainfrom
mordante:compact_ucd_tables
Jun 12, 2022
Merged

Stephan T. Lavavej (StephanTLavavej) merged 2 commits into
microsoft:mainfrom
mordante:compact_ucd_tables

Conversation

@mordante

Copy link
Copy Markdown
Contributor

The Unicode data files used to generate the extended grapheme cluster
tables contain contiguous ranges split over several lines, for example:

2060..2064 ; Control # Cf [5] WORD JOINER..INVISIBLE PLUS
2065 ; Control # Cn
2066..206F ; Control # Cf [10] LEFT-TO-RIGHT ISOLATE..NOMINAL DIGIT SHAPES

Instead of creating separate table entries, create one entry for the
combined range. Especially the Extended_Pictographic property has a lot
of ranges that can be combined. Combining these ranges reduces the size
of the tables and should improve performance.

This change reduces the total number of entries in the tables
_Grapheme_Break_property_data and _Extended_Pictographic_property_values
by 445, saving 2670 bytes of data.

The Unicode data files used to generate the extended grapheme cluster
tables contain contiguous ranges split over several lines, for example:

  2060..2064    ; Control # Cf   [5] WORD JOINER..INVISIBLE PLUS
  2065          ; Control # Cn       
  2066..206F    ; Control # Cf  [10] LEFT-TO-RIGHT ISOLATE..NOMINAL DIGIT SHAPES

Instead of creating separate table entries, create one entry for the
combined range. Especially the Extended_Pictographic property has a lot
of ranges that can be combined. Combining these ranges reduces the size
of the tables and should improve performance.

This change reduces the total number of entries in the tables
_Grapheme_Break_property_data and _Extended_Pictographic_property_values
by 445, saving 2670 bytes of data.
@mordante
Mark de Wever (mordante) requested a review from a team as a code owner June 3, 2022 18:28
@mordante

Copy link
Copy Markdown
Contributor Author

Charlie Barto (@barcharcraz) I assume you want to have a look at this.

Comment thread tools/unicode_properties_parse/grapheme_break_property_data_gen.py Outdated
Comment thread tools/unicode_properties_parse/grapheme_break_property_data_gen.py Outdated
@StephanTLavavej

Copy link
Copy Markdown
Member

Thanks for noticing this and writing an elegant improvement! 😻 I've pushed very small changes to the Python script for the function name and comment.

Additionally, I have verified that the output of the script, followed by clang-format, exactly matches the change to the product header.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is a great change! It looks like they repeat stuff like this in order to keep things more "stable" over unicode updates, and to split "unrelated" emoji, even if they happen to be packed in one after the other.

We don't need to worry about either of those things here.

@barcharcraz

Copy link
Copy Markdown
Contributor

Oh, I think I'm wrong about why they do this. It's to line up the ranges with other properties! So for Gbpdata if a character has the same GBP as the preceding range but a different General category things are split. They must figure it's easier to rejoin the ranges than to split them apart, which is quite true.

For emoji-data it's more common to see the breaks, because they are breaking the ranges if any of the other emoji properties are different.

@StephanTLavavej

Copy link
Copy Markdown
Member

I'm mirroring this to the MSVC-internal repo - please notify me if any further changes are pushed.

@StephanTLavavej
Stephan T. Lavavej (StephanTLavavej) merged commit 82acfcf into microsoft:main Jun 12, 2022
@StephanTLavavej

Copy link
Copy Markdown
Member

Thx 4 tiny tbls! 😹 📉 🎉

@CaseyCarter

Casey Carter (CaseyCarter) commented Jun 12, 2022 •

Copy link
Copy Markdown
Contributor

Congratulations on stepping up to contribute to the best open-source C++ Standard Library! :trollface:

@mordante

Copy link
Copy Markdown
Contributor Author

Congratulations on stepping up to contribute to the best open-source C++ Standard Library! :trollface:

Thanks for acknowledging my contributions to libc++ :-P

@CaseyCarter

Copy link
Copy Markdown
Contributor

Congratulations on stepping up to contribute to the best open-source C++ Standard Library! :trollface:

Thanks for acknowledging my contributions to libc++ :-P

🔥🔥🔥

Igor Zhukov (fsb4000) pushed a commit to fsb4000/STL that referenced this pull request Aug 13, 2022
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

format C++20/23 format performance Must go faster

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants