Skip to content

Add the read_article event tracking #13287

Description

@eugene-manuilov

Feature Description

GA4 Enhanced Measurement's scroll event fires when a visitor reaches 90% of the page, on every page type — the home page, archives, static pages, custom post types. It cannot answer "did someone read this article", and it counts a visitor who jumps straight to the bottom.

This issue adds read_article. It is sent only on a single blog post, and only when the visitor both reaches the end of the article and stays long enough to have read it. The work has two halves.

1. Server side — the Content_Events provider (#13281) appends an invisible marker to the end of the post content through the the_content filter, and measures the length of that same content. the_content runs many times per request, and the first run is not reliably the main content, so the filter is guarded case by case: a nested loop rendering another post, an automatic excerpt of the current post, a feed request, the oEmbed template, and every page of a paginated post except the last. The provider then adds the word count, the estimated reading time, the reading-time constants and an isFinalPage flag to the config object introduced in #13281.

2. Frontend — content-events.ts (#13281) gains the read_article handler: an IntersectionObserver watching the marker, a scroll-depth fallback for page builders that never run the_content, and a timer that only runs while the tab is visible and focused. Both halves of the trigger are gated on the isSinglePost and isFinalPage flags the server publishes.

How the reading time is estimated. The estimate is the word count divided by 238 words per minute, and the part that breaks outside English is the word count. Chinese, Japanese, Thai, Khmer, Lao and Burmese are written without spaces between words, so splitting the text on spaces returns one word per paragraph: a 1,120-character Japanese post counts as one word, its estimated reading time falls below a second, the 5-second floor takes over, and the time half of the trigger disappears. The provider therefore counts words with ICU's dictionary-based segmenter, IntlBreakIterator, which knows those scripts — 我喜欢写代码。 counts as 4 words, 私はコードを書くのが好きです。 as 9 — and returns exactly the same count as a space split for every language that separates words with spaces, so nothing changes for content that is counted correctly today. The result does not depend on the site locale, so a multilingual site needs no language detection. intl is an optional PHP extension, so where it is missing the provider falls back to counting words by spaces plus the characters of those six scripts at 500 characters per minute.

Paginated posts. On a post split with the marker is added to the last page only, so the post sends at most one read_article per visit instead of one per page. This is why the frontend cannot read "no marker" as "the content pipeline was bypassed", and needs isFinalPage from the server: on the earlier pages the marker is missing on purpose, and an ungated scroll fallback would send one event per page — exactly what the flag prevents. A visitor who reads pages 1 to 3 of 4 and leaves sends nothing.

Link to the design doc: https://docs.google.com/document/d/1F23KG9do9PMzZpe-KhCfGw97C5BFIrEt5otC4j4ITn0/edit?tab=t.saezzs38gdl


Do not alter or remove anything below. The following sections will be managed by moderators only.

Acceptance criteria

  • On a single blog post, one read_article event is sent to GA once both of these have happened, in either order:
    • the end of the article text has been on screen;
    • the visitor has spent 85% of the article's estimated reading time on the page, and never less than 5 seconds.
  • The event carries post_id, word_count and estimated_read_time_seconds, and nothing else.
  • Time while the browser tab is hidden or the window is in the background does not count towards the waiting time. It resumes from where it stopped when the visitor comes back.
  • The estimated reading time is the article's word count divided by 238 words per minute:
    The article contains Estimated reading time The visitor must stay
    238 words 60s 51s
    476 words 120s 102s
    20 words 5s 5s — the floor, not 4.25s
    a Japanese post of about 1,120 characters, which is 600 words 151s 128s
  • Words are counted the way each writing system writes them, so a language that does not separate words with spaces is still counted correctly:
    The article text word_count
    The quick brown fox jumps over the lazy dog. 9
    我喜欢写代码。 4 — 我 / 喜欢 / 写 / 代码
    私はコードを書くのが好きです。 9 — 私 / は / コード / を / 書く / の / が / 好き / です
    Site Kit by Google の新機能 7 — the four English words and the three Japanese ones
    • English, Korean, Arabic, Russian, Greek, Hebrew, Hindi and every other language that separates words with spaces is counted exactly as it is today.
    • The count does not change with the site's language setting: a Japanese post counts the same on a site running in English.
    • Digits count as words. Punctuation, emoji and HTML markup do not, and shortcodes are removed before counting: Version 3.14 is 2 words, 「テスト」、。!? is 1, and Hello 👋 world is 2.
  • On a server without the PHP intl extension, read_article still fires. Words are counted by spaces there, and every Chinese, Japanese, Thai, Lao, Khmer and Burmese character is counted at 500 characters per minute, so a post in those languages still has to be read rather than only reached.
  • One page view sends at most one event. Scrolling away from the end of the article and back, or staying on the page longer, sends nothing more.
  • Nothing is sent on the home page, an archive, a search results page, a static page or a single custom post type page, however far the visitor scrolls and however long they stay.
  • The rendered post content ends with an invisible marker: an HTML comment. Hidden from screen readers, that a visitor cannot see or interact with. It sits directly after the author's content, before anything other plugins append to the end of a post — share buttons, related posts, subscribe forms.
  • The marker appears generally once per page, and never:
    • in an RSS feed;
    • in a post excerpt, including an excerpt of the post currently being viewed;
    • in the content of another post shown on the page by a Query Loop or a related-posts block;
    • on the WordPress oEmbed page for the post;
    • on a page, a custom post type, an archive or the home page;
    • on any page of a post other than the last.
  • BUT: A theme that prints the post twice/multiple times can show the marker in both copies.
  • On a post, pages 1 to n−1 send no event however far the visitor scrolls and however long they stay. The last page behaves like any other post, and its word count covers that page's text only.
  • When the page source contains no marker — a page builder that produces the post content without running the_content — scrolling to within 10% of the bottom of the page counts as reaching the end of the article instead.

Implementation Brief

  • In includes/Core/Conversion_Tracking/Conversion_Event_Providers/Content_Events.php (added in Content_Events provider scaffold #13281):

    • Add the reading-time constants block: const WORDS_PER_MINUTE = 238;, const READ_TIME_THRESHOLD_PERCENT = 85;, const MINIMUM_READ_TIME_SECONDS = 5;, and const FALLBACK_CHARACTERS_PER_MINUTE = 500; used only by the path below that runs without intl.
    • Add const END_OF_CONTENT_ANCHOR = '';.
    • Add const SCRIPTS_WITHOUT_WORD_SPACING = '\p{Han}\p{Hiragana}\p{Katakana}\p{Thai}\p{Lao}\p{Khmer}\p{Myanmar}'; — the character-class body the fallback counts with.
    • Add protected $content_captured = false;, protected $word_count = null;, protected $estimated_read_time_seconds = null; and protected $is_final_page = null;.
    • In register_content_hooks(), add a the_content filter at priority 1 calling a new protected function append_end_of_content_anchor( $content ).
    • append_end_of_content_anchor():
      • Return $content unchanged when any of these hold: $this->content_captured; ! is_singular( 'post' ); is_feed(); is_embed(); doing_filter( 'get_the_excerpt' ); get_the_ID() !== get_queried_object_id().
      • Otherwise set $this->content_captured = true;, then, from the $page / $numpages / $multipage globals, set $this->is_final_page = ! $multipage || $page >= $numpages; and fill $this->word_count / $this->estimated_read_time_seconds from $this->measure_content( $content ).
      • Return $content . self::END_OF_CONTENT_ANCHOR when $this->is_final_page, and $content unchanged otherwise.
    • Add protected function measure_content( $content ) returning array( 'word_count' => int, 'estimated_read_time_seconds' => int ):
      • Strip first: $text = wp_strip_all_tags( strip_shortcodes( $content ) );.
      • Take $word_count from $this->count_words_with_intl( $text ). When it is not null, the estimate is (int) round( $word_count / self::WORDS_PER_MINUTE * 60 ).
      • When it is null, take $word_count from $this->count_words_by_spaces( $text ) and the estimate from (int) round( ( $word_count / self::WORDS_PER_MINUTE + $this->count_characters_without_word_spacing( $text ) / self::FALLBACK_CHARACTERS_PER_MINUTE ) * 60 ).
    • Add protected function count_words_with_intl( $text ) returning int|null — this is the seam the fallback tests replace, so keep it a separate method:
      • Return null when ! class_exists( 'IntlBreakIterator' ).
      • $iterator = IntlBreakIterator::createWordInstance( get_locale() ); — return null when it is falsy or when $iterator->setText( $text ) returns false. The locale only picks ICU's tailoring; the dictionaries for Han, Kana, Thai, Khmer, Lao and Burmese are used whatever it is, and an unknown WordPress locale such as de_DE_formal falls back inside ICU rather than failing.
      • Iterate $iterator->getPartsIterator() and count the parts matching /[\p{L}\p{N}]/u, so spaces and punctuation are skipped and numbers are kept.
    • Add protected function count_words_by_spaces( $text ) — preg_split( '/\s+/u', trim( $text ), -1, PREG_SPLIT_NO_EMPTY ), counting only the pieces that match /[\p{L}\p{N}]/u. Return 0 when preg_split() returns false, which is what it does for content that is not valid UTF-8; count( false ) is a fatal error on PHP 8.
    • Add protected function count_characters_without_word_spacing( $text ) — preg_match_all( '/(?=[' . self::SCRIPTS_WITHOUT_WORD_SPACING . '])[\p{L}\p{N}\p{M}]/u', $text ), returning 0 when it returns false. The lookahead plus the letter/digit/mark class is what keeps CJK punctuation out of the count: PCRE resolves \p{Han} through Unicode Script Extensions, so 、, 。, 「 and 」 match the script class on their own.
    • Extend the array returned by get_inline_config() with:
      • wordCount and estimatedReadTimeSeconds — (int) casts of the captured values; when the filter never ran, measure_content( get_the_content( null, false, get_queried_object_id() ) ) on a single post and 0 for both otherwise.
      • isFinalPage — (bool) $this->is_final_page; when the filter never ran, derive it on a single post from generate_postdata( get_queried_object_id() ) as ! $multipage || $page >= $numpages, and use false otherwise.
      • readTimeThresholdPercent and minimumReadTimeSeconds from the constants above. The rates stay in PHP; the frontend never sees them.
  • In assets/js/event-providers/content-events/constants.ts (new file):

    • Export READ_TIME_THRESHOLD_PERCENT = 85 and MINIMUM_READ_TIME_SECONDS = 5, matching the PHP constants of the same name — they are the defaults used when the config does not carry them.
    • Export END_OF_CONTENT_SELECTOR = '.googlesitekit-end-of-content' and SCROLL_DEPTH_THRESHOLD = 0.9.
  • In assets/js/event-providers/content-events/read-article.ts (new file):

    • Export initializeReadArticle( config: ContentEventsConfig ): void, taking the config type content-events.ts exports; add that exported type there if the TypeScript conversion of Content_Events provider scaffold #13281 did not.
    • Return immediately unless config.isSinglePost && config.isFinalPage — registering no observer, no scroll listener and no timer in that case.
    • Required dwell in milliseconds: Math.max( minimumReadTimeSeconds, ( estimatedReadTimeSeconds * readTimeThresholdPercent ) / 100 ) * 1000.
    • Position condition:
      • When document.querySelector( END_OF_CONTENT_SELECTOR ) finds an element, observe it with an IntersectionObserver; mark the condition met on the first entry with isIntersecting, then disconnect().
      • Otherwise add a passive scroll listener on global and mark the condition met once ( global.scrollY + global.innerHeight ) / document.documentElement.scrollHeight >= SCROLL_DEPTH_THRESHOLD, then remove the listener. Evaluate the same check once on init.
    • Dwell condition — accumulate only visible, focused time:
      • On init, and whenever the document becomes visible and focused again, start a setTimeout for the time still outstanding.
      • On visibilitychange to hidden and on blur, clear the timeout and add the elapsed slice to the accumulated total.
      • Register the visibilitychange, blur and focus listeners only when the handler is active.
    • When both conditions have been met, in either order, emit exactly once through global._googlesitekit?.gtagEvent?.( 'read_article', { post_id: postID, word_count: wordCount, estimated_read_time_seconds: estimatedReadTimeSeconds } ), then disconnect the observer, clear the timeout and remove every listener the handler added.
  • In assets/js/event-providers/content-events.ts (added in Content_Events provider scaffold #13281):

    • Extend the getContentEventsConfig() defaults and its exported type with wordCount: 0, estimatedReadTimeSeconds: 0, isFinalPage: false, and readTimeThresholdPercent / minimumReadTimeSeconds from content-events/constants.ts.
    • Call initializeReadArticle( getContentEventsConfig() ) from a try/catch that swallows the error, so a throw cannot stop the handlers added by the follow-up issues.

Test Coverage

  • Extend tests/phpunit/integration/Core/Conversion_Tracking/Conversion_Event_Providers/Content_EventsTest.php (added in Content_Events provider scaffold #13281) covering:

    • The anchor is appended to the queried post's content once the content hooks are bootstrapped, and its markup is the 1px, aria-hidden, block-level .
    • The anchor precedes markup appended by another the_content filter registered at priority 10.
    • A second application of the_content in the same request does not append a second anchor.
    • No anchor when the filter runs for a different post (global $post swapped with setup_postdata()), when get_the_excerpt() is applied for the current post both before and after the real content, on a feed request, on the oEmbed template, and on a page, a CPT single, an archive and the home page.
    • On a post: no anchor on pages 1…n−1, anchor on the last page.
    • The word count covers the filtered content with shortcodes, tags and block delimiter comments stripped, and covers only the current page's text on a paginated post.
    • A data provider over writing systems, skipped with markTestSkipped() when IntlBreakIterator is missing on the runner: 我喜欢写代码。 is 4 words; 私はコードを書くのが好きです。 is 9; Site Kit by Google の新機能 is 7; a Thai sentence counts more than one word; and English, Korean, Arabic, Russian, Greek, Hebrew and Hindi fixtures each count the same as their space-separated word count.
    • The same Japanese fixture returns the same count with get_locale() filtered to en_US, ja_JP and zh_CN.
    • Version 3.14 is 2 words, 「テスト」、。!? is 1, Hello 👋 world is 2.
    • estimated_read_time_seconds is 60 for a 238-word post and 120 for a 476-word one.
    • The fallback path, exercised through a subclass whose count_words_with_intl() returns null: an English fixture keeps its word count and its estimate, and a Japanese fixture is estimated from its characters at 500 per minute instead of falling to the floor.
    • Content that is not valid UTF-8 raises no PHP warning or error: the ICU path skips the invalid part and counts the rest of the text, and the fallback path returns 0.
    • The inline config carries wordCount, estimatedReadTimeSeconds, readTimeThresholdPercent, minimumReadTimeSeconds and isFinalPage, with isFinalPage true for an unpaginated post and for the last page of a paginated one, and false for earlier pages.
    • The word count, the estimate and isFinalPage fall back to values derived from the queried post when the content hooks are bootstrapped but the_content never runs.
  • Add assets/js/event-providers/content-events/read-article.test.ts, mocking IntersectionObserver with intersectionObserver from @shopify/jest-dom-mocks (as assets/js/hooks/useLatestIntersection.test.js does) and using Jest fake timers, covering:

    • Nothing is registered — no observer, no scroll listener, no timer — when isSinglePost or isFinalPage is false.
    • The event fires once with post_id, word_count and estimated_read_time_seconds when the anchor intersects and the dwell threshold is then reached, and when the two happen in the opposite order.
    • Neither condition alone fires the event.
    • The required dwell is 85% of estimatedReadTimeSeconds: 120s requires 102s, 151s requires 128.35s, and the event does not fire one second earlier in either case.
    • An estimatedReadTimeSeconds of 5 or less uses the minimumReadTimeSeconds floor.
    • The scroll fallback is used when the anchor is absent, firing at ≥90% depth and not before, and a page already scrolled to the bottom on init satisfies the position condition.
    • Time spent while the document is hidden or the window is blurred does not count towards the dwell, and the timer resumes on focus.
    • Further intersections, scrolls or timer ticks after the event has fired emit nothing.
  • Extend assets/js/event-providers/content-events.test.ts (added in Content_Events provider scaffold #13281) with:

    • The config helper's defaults for wordCount, estimatedReadTimeSeconds, isFinalPage and the two threshold keys, and the published values when the config supplies them.
    • A handler that throws does not propagate the error out of the entry module.

QA Brief

Before you start

  • Connect Analytics 4, and turn Place Google Analytics code and Plugin conversion tracking on.
  • Keep the Network panel open and filtered on collect. Every step that follows reads the rows the Network panel lists.
  • Install a plugin that appends related posts or share buttons to the end of a post. Test 1 uses it.
  • Publish the eight posts in the table that follows. You must also have one static page and one single post of a custom post type.
Post Its content
English 238 words.
Digits and emoji Version 3.14 Hello 👋 world [gallery] and nothing else.
Chinese 我喜欢写代码。
Japanese and English Site Kit by Google の新機能
Long Japanese コードを書く。 repeated 200 times, which Site Kit counts as 600 words.
Paginated Two pages, split with . Each page has 238 words.
Page builder 20 words, rendered by a page builder rather than by the_content.
Vimeo A Vimeo video added as an embed block.

Test 1: A post with other content below its text

  1. Open the English post. Confirm that no visible gap sits between the end of the post text and whatever the theme renders next.
  2. Add a related-posts block or a Query Loop block below the text of the English post, then open the post. Scroll to the bottom and wait 60 seconds. Confirm that one read_article row appears, and no more than one.
  3. Activate the plugin that appends related posts or share buttons, then open the English post. Scroll until the post text ends, then stop and wait 60 seconds. Confirm that one read_article row appears before you scroll down again.

Test 2: The two halves of the trigger

The English post is estimated at 60 seconds, so a visitor must get to the end of its text and stay 51 seconds.

  1. Open the English post, scroll to the bottom, and wait 20 seconds. Confirm that a scroll row appears, and that no read_article row appears.
  2. Keep that page open for another 40 seconds. Confirm that one read_article row appears, and that its query string has en=read_article. Confirm that post_id, word_count, and estimated_read_time_seconds are the only three parameters the event adds.
  3. Scroll that page to the top and back to the bottom three times. Confirm that the read_article row count stays at one.
  4. Reload the English post, then stay at the top for 90 seconds without scrolling. Confirm that no read_article row appears.
  5. Scroll that same page to the bottom. Confirm that a read_article row appears in a few seconds.

Test 3: Time while the tab is hidden or the window sits behind another

  1. Reload the English post, scroll to the bottom, then switch to another browser tab for a minute. Confirm that no read_article row appears while you are away. Confirm that one appears about 50 seconds after you come back.
  2. Reload the English post, scroll to the bottom, then click a window from another program for a minute. Confirm that no read_article row appears while Chrome sits behind that window. Confirm that one appears about 50 seconds after you click back.

Test 4: The word count in each writing system

Run each post in this table on its own. Open it, scroll to the bottom, wait the number of seconds the last column gives, then read the read_article row.

Post word_count estimated_read_time_seconds Wait, in seconds
Digits and emoji 4 1 5
Chinese 4 1 5
Japanese and English 7 2 5
Long Japanese 600 151 150
  1. Confirm that each post sends one read_article row, with the word_count and the estimated_read_time_seconds its row gives.
  2. In Settings > General, set Site Language to Japanese. Reload the long Japanese post, scroll to the bottom, and wait 150 seconds. Confirm that word_count still reads 600.

Test 5: A server without the PHP intl extension

  1. Turn the PHP intl extension off, then restart PHP.
  2. Open the long Japanese post, scroll to the bottom, and wait 150 seconds. Confirm that one read_article row appears, with word_count at 1 and estimated_read_time_seconds at 144.
  3. Turn the intl extension back on, then restart PHP again.

Test 6: A post split into pages

  1. Open page 1 of the paginated post, scroll to the bottom, and wait a minute. Confirm that no read_article row appears.
  2. Open page 2 of the paginated post, scroll to the bottom, and wait 60 seconds. Confirm that one read_article row appears, and that its word_count reads 238 rather than 476.

Test 7: The pages that send no read_article event

  1. Open the home page, the static page, the single post of the custom post type, a category archive, and a search results page in turn. Scroll each one to the bottom and wait 10 seconds. Confirm that a scroll row appears on each one. Confirm that no read_article row appears on any of them.

Test 8: A post a page builder renders

  1. Open the page builder post, scroll to the bottom of the page, and wait five seconds. Confirm that one read_article row appears, and that the markup has no googlesitekit-end-of-content.

Test 9: The pagination and Vimeo events still fire

  1. Open the paginated post, then click its page 2 link. Confirm that a pagination_click row appears.
  2. Open the Vimeo post, then play the video. Confirm that a video_start row appears.

Note: the built-in scroll event in Analytics 4 still fires at 90% of every page, read_article included. Two rows in the Network panel on one post view are expected.

Changelog entry

  • Add event tracking that monitors when a reader finishes reading a post.

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    P1Medium priorityTeam SIssues for Squad 1Type: EnhancementImprovement of an existing feature

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions