You are currently viewing Text Encoding Basics Before You Convert

Text Encoding Basics Before You Convert

TL;DR: Text encoding is the agreement about which bytes mean which characters. UTF-8, legacy Chinese encodings such as GBK, and HTML entities solve different problems. Convert only when you know both the encoding you have and the encoding the destination expects.

Bytes are not the letters you see

A text file is a sequence of bytes plus an assumption about how to read them. UTF-8 is the assumption most of the web uses now. Older Chinese pages and some desktop software still use GBK or GB2312. If you open GBK bytes as if they were UTF-8, you do not get a mildly misspelled sentence. You get replacement characters, question marks, or a string of Latin letters that never formed a word. The file did not “get corrupted” in a vague sense. It was interpreted with the wrong map.

Masters Tool’s reference tables, including ASCII and GBK listings, are maps you consult. They do not change your file. The encoding conversion tools, such as HTML entity encoding, change a string from one written form to another. Use a table when you are asking “what is this code.” Use a converter when you are asking “rewrite this string so the next system accepts it.”

ASCII, entities, and the characters beyond them

ASCII covers a small set: English letters, digits, and basic punctuation. The moment a sentence includes an accent, a curly quote, a degree sign, or any Chinese character, you are outside ASCII. HTML entities exist so a character that would break a markup file can be written with plain ASCII anyway, like a named or numeric reference. That is a convenience for HTML. It is not a general encoding for storing documents. If you entity-encode a paragraph and paste it into a place that expected raw UTF-8, the reader will see the entity codes instead of the letters.

Curly quotes and non-breaking spaces are the everyday version of this pain. A word processor inserts them. A strict code tool or a CSV parser chokes, or it splits a field in a place you did not intend. Converting those characters to plain quotes and ordinary spaces is a reasonable cleaning step. Converting an entire multilingual paragraph to ASCII by dropping “unsupported” letters is not cleaning. It deletes the language.

When a conversion is actually called for

Convert when a destination has told you its requirement. A stylesheet or an HTML attribute might need a URL-encoded or entity-encoded fragment. A legacy importer might demand GBK. A programming literal might need escape sequences. In each case you should be able to name the source form and the target form. “Make this work” is not a target. If you do not know the target, look at a file that destination already accepts, and match that, or read the error. Guessing UTF-8 when the importer wanted a legacy code page will fail in a way that looks like random punctuation.

Do not convert twice. Entity-encoding a string that is already entity-encoded turns an ampersand into a second layer of codes, and the page will display the codes rather than the character. Base64, which belongs to binary-as-text rather than to character sets, has the same double-wrap failure: encoding an already encoded blob produces a longer string that decodes to the previous string, not to the original file. If you are unsure, decode once in a scratch note and see whether the result is readable text or a sensible file header.

Separators and sample lines

  • Open a known word. If a Chinese greeting or a simple accented word round-trips, the pair of encodings is plausible. If it does not, stop and change the assumption.
  • Keep a raw copy of the original bytes before any conversion. The converted file is not a backup.
  • Watch the BOM. A byte-order mark at the start of UTF-8 is invisible in some editors and visible as junk in others. It can break a script’s first line.
  • Newlines are not encodings, but they travel with the file. A tool that also rewrites line endings can confuse a later diff. Know whether you asked for that.
  • HTML source is not the rendered page. Copying from a rendered page gives you characters. Copying from source may give you entities. Convert only the one you meant to copy.

Tables help you verify a single character

When one character survives badly, do not re-encode the whole document in a loop. Look that character up. An ASCII table tells you the low codes. A GBK table tells you whether the character even exists in that set. If it does not exist in the target set, no converter can represent it faithfully. You must pick a different target, or accept a replacement that you choose on purpose, such as a romanization or a description, rather than a silent question mark.

Programmers hit this when a config file, a CSV, and a database column disagree. The practical fix is to agree on UTF-8 at every boundary you control, and to convert only at the boundary you do not control. Scattershot conversion at each step is how a name becomes unreadable by the third system.

A cautious first conversion

Take a ten-character sample that includes one non-ASCII letter you can recognize. Run it through the encoder you think you need. Confirm the sample by eye. Only then run the rest. If the tool offers a charset dropdown, the dropdown is the whole decision. Leaving it on the default is a decision too, and it may not match your file.

Reach those encoders and tables from the Masters Tool homepage after you can say the source form and the target form out loud. If you cannot say both, you are not ready to press the button.