KITFORMA · EDITORIAL ANSWER
Why do UTF-8 characters become corrupted when decoding network chunks separately?
KitForma editorial guide. Network chunk boundaries do not align with character boundaries. A multi-byte character can be split across two chunks.
Step-by-step answer
import assert from 'node:assert/strict';
import { StringDecoder } from 'node:string_decoder';
const bytes = Buffer.from('€', 'utf8');
for (let split = 1; split < bytes.length; split++) {
const decoder = new StringDecoder('utf8');
const text = decoder.write(bytes.subarray(0, split))
+ decoder.write(bytes.subarray(split)) + decoder.end();
assert.equal(text, '€');
}Sources and verification
Sources checked:
Scope: This editorial guide is based on the cited sources and tool behavior. A forum question or a query observed for our site does not establish market search volume, low competition, guaranteed rankings or inadequate answers elsewhere.
This editorial answer was prepared by KitForma with AI assistance. It is not presented as a real member question or an independent user review. Check the sources and the result with your own file; report corrections in the discussion.