{"id":50,"date":"2026-10-07T17:11:00","date_gmt":"2026-10-07T17:11:00","guid":{"rendered":"https:\/\/www.masters-tool.com\/blog\/text-encoding-basics-before-you-convert\/"},"modified":"2026-10-08T03:04:42","modified_gmt":"2026-10-08T03:04:42","slug":"text-encoding-basics-before-you-convert","status":"publish","type":"post","link":"https:\/\/www.masters-tool.com\/blog\/text-encoding-basics-before-you-convert\/","title":{"rendered":"Text Encoding Basics Before You Convert"},"content":{"rendered":"<p><strong>TL;DR:<\/strong> Text encoding is the agreement about which bytes mean which characters. UTF-8, legacy Chinese encodings such as GBK, and HTML entities solve different problems. Convert only when you know both the encoding you have and the encoding the destination expects.<\/p>\n<h2>Bytes are not the letters you see<\/h2>\n<p>A text file is a sequence of bytes plus an assumption about how to read them. UTF-8 is the assumption most of the web uses now. Older Chinese pages and some desktop software still use GBK or GB2312. If you open GBK bytes as if they were UTF-8, you do not get a mildly misspelled sentence. You get replacement characters, question marks, or a string of Latin letters that never formed a word. The file did not \u201cget corrupted\u201d in a vague sense. It was interpreted with the wrong map.<\/p>\n<p>Masters Tool\u2019s reference tables, including ASCII and GBK listings, are maps you consult. They do not change your file. The encoding conversion tools, such as HTML entity encoding, change a string from one written form to another. Use a table when you are asking \u201cwhat is this code.\u201d Use a converter when you are asking \u201crewrite this string so the next system accepts it.\u201d<\/p>\n<h2>ASCII, entities, and the characters beyond them<\/h2>\n<p>ASCII covers a small set: English letters, digits, and basic punctuation. The moment a sentence includes an accent, a curly quote, a degree sign, or any Chinese character, you are outside ASCII. HTML entities exist so a character that would break a markup file can be written with plain ASCII anyway, like a named or numeric reference. That is a convenience for HTML. It is not a general encoding for storing documents. If you entity-encode a paragraph and paste it into a place that expected raw UTF-8, the reader will see the entity codes instead of the letters.<\/p>\n<p>Curly quotes and non-breaking spaces are the everyday version of this pain. A word processor inserts them. A strict code tool or a CSV parser chokes, or it splits a field in a place you did not intend. Converting those characters to plain quotes and ordinary spaces is a reasonable cleaning step. Converting an entire multilingual paragraph to ASCII by dropping \u201cunsupported\u201d letters is not cleaning. It deletes the language.<\/p>\n<h2>When a conversion is actually called for<\/h2>\n<p>Convert when a destination has told you its requirement. A stylesheet or an HTML attribute might need a URL-encoded or entity-encoded fragment. A legacy importer might demand GBK. A programming literal might need escape sequences. In each case you should be able to name the source form and the target form. \u201cMake this work\u201d is not a target. If you do not know the target, look at a file that destination already accepts, and match that, or read the error. Guessing UTF-8 when the importer wanted a legacy code page will fail in a way that looks like random punctuation.<\/p>\n<p>Do not convert twice. Entity-encoding a string that is already entity-encoded turns an ampersand into a second layer of codes, and the page will display the codes rather than the character. Base64, which belongs to binary-as-text rather than to character sets, has the same double-wrap failure: encoding an already encoded blob produces a longer string that decodes to the previous string, not to the original file. If you are unsure, decode once in a scratch note and see whether the result is readable text or a sensible file header.<\/p>\n<h2>Separators and sample lines<\/h2>\n<ul>\n<li><strong>Open a known word.<\/strong> If a Chinese greeting or a simple accented word round-trips, the pair of encodings is plausible. If it does not, stop and change the assumption.<\/li>\n<li><strong>Keep a raw copy<\/strong> of the original bytes before any conversion. The converted file is not a backup.<\/li>\n<li><strong>Watch the BOM.<\/strong> A byte-order mark at the start of UTF-8 is invisible in some editors and visible as junk in others. It can break a script\u2019s first line.<\/li>\n<li><strong>Newlines are not encodings,<\/strong> but they travel with the file. A tool that also rewrites line endings can confuse a later diff. Know whether you asked for that.<\/li>\n<li><strong>HTML source is not the rendered page.<\/strong> Copying from a rendered page gives you characters. Copying from source may give you entities. Convert only the one you meant to copy.<\/li>\n<\/ul>\n<h2>Tables help you verify a single character<\/h2>\n<p>When one character survives badly, do not re-encode the whole document in a loop. Look that character up. An ASCII table tells you the low codes. A GBK table tells you whether the character even exists in that set. If it does not exist in the target set, no converter can represent it faithfully. You must pick a different target, or accept a replacement that you choose on purpose, such as a romanization or a description, rather than a silent question mark.<\/p>\n<p>Programmers hit this when a config file, a CSV, and a database column disagree. The practical fix is to agree on UTF-8 at every boundary you control, and to convert only at the boundary you do not control. Scattershot conversion at each step is how a name becomes unreadable by the third system.<\/p>\n<h2>A cautious first conversion<\/h2>\n<p>Take a ten-character sample that includes one non-ASCII letter you can recognize. Run it through the encoder you think you need. Confirm the sample by eye. Only then run the rest. If the tool offers a charset dropdown, the dropdown is the whole decision. Leaving it on the default is a decision too, and it may not match your file.<\/p>\n<p>Reach those encoders and tables from the <a href=\"https:\/\/www.masters-tool.com\/\">Masters Tool homepage<\/a> after you can say the source form and the target form out loud. If you cannot say both, you are not ready to press the button.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Name the encoding you have and the one the destination expects. Do not encode a string that is already encoded.<\/p>\n","protected":false},"author":1,"featured_media":39,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[3],"tags":[],"class_list":["post-50","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-converters-codecs","entry","has-media"],"_links":{"self":[{"href":"https:\/\/www.masters-tool.com\/blog\/wp-json\/wp\/v2\/posts\/50","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.masters-tool.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.masters-tool.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.masters-tool.com\/blog\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.masters-tool.com\/blog\/wp-json\/wp\/v2\/comments?post=50"}],"version-history":[{"count":1,"href":"https:\/\/www.masters-tool.com\/blog\/wp-json\/wp\/v2\/posts\/50\/revisions"}],"predecessor-version":[{"id":51,"href":"https:\/\/www.masters-tool.com\/blog\/wp-json\/wp\/v2\/posts\/50\/revisions\/51"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.masters-tool.com\/blog\/wp-json\/wp\/v2\/media\/39"}],"wp:attachment":[{"href":"https:\/\/www.masters-tool.com\/blog\/wp-json\/wp\/v2\/media?parent=50"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.masters-tool.com\/blog\/wp-json\/wp\/v2\/categories?post=50"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.masters-tool.com\/blog\/wp-json\/wp\/v2\/tags?post=50"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}