encodeURIComponent vs encodeURI, and where %2520 and "URI malformed" come from
Percent-encoding looks like one operation, but there are three common variants, and mixing them up
produces the same handful of bugs: a query that loses half its parameters, a %2520 in the address bar,
a URIError, and Chinese text turning into question marks. Each has one cause.
Which characters each function leaves alone
encodeURIComponent leaves only letters, digits and - _ . ! ~ * ' ( ). encodeURI leaves those plus
the characters that structure an address: ; , / ? : @ & = + $ #.
| Character | encodeURIComponent | encodeURI |
|---|---|---|
/ | %2F | kept |
? | %3F | kept |
# | %23 | kept |
& | %26 | kept |
= | %3D | kept |
+ | %2B | kept |
% | %25 | %25 |
| space | %20 | %20 |
中 | %E4%B8%AD | %E4%B8%AD |
encodeURIComponent("a b&c=中"); // "a%20b%26c%3D%E4%B8%AD"
encodeURI("a b&c=中"); // "a%20b&c=%E4%B8%AD"
The rule that follows: encodeURIComponent is for one piece that will sit inside a URL, such as a query
value or a path segment. encodeURI is for an address that is already assembled, where you only want
spaces and non-ASCII characters escaped and the separators left working.
A value with & or = splits the query
The symptom is a parameter that arrives truncated, or an extra parameter nobody sent. It appears when a
value is built with encodeURI, or with no encoding at all:
const q = "R&D = 50% off";
"/search?q=" + encodeURI(q); // /search?q=R&D%20=%2050%25%20off
"/search?q=" + encodeURIComponent(q); // /search?q=R%26D%20%3D%2050%25%20off
The first form is read as q=R plus a second parameter named D . The second keeps the whole phrase in
q. encodeURI cannot help here because it has no way to know that the & is data rather than a
separator. Encode each value separately, then join with literal & and =. Or skip the manual work:
new URLSearchParams({ q }).toString() does the escaping for you.
The same applies to #. An unescaped # in a value starts the fragment, and everything after it is
never sent to the server.
Space is %20 in one place and + in another
%20 is correct everywhere in a URL. + means a space only inside a query string or a
application/x-www-form-urlencoded body, which is what an HTML form submits. In a path, + is a literal
plus sign. Decoding shows the difference:
decodeURIComponent("a+b"); // "a+b"
new URLSearchParams("a=1+2").get("a"); // "1 2"
Two bugs come from this. A value containing a real + (a phone number, c++, a base64 string) loses it
when it is encoded with a form-style encoder that does not escape it, because the receiver turns it into a
space; it has to travel as %2B. And a + that ends up in a path stays a plus, so a “space” you meant
there is wrong. In the URL encoder, the Form method turns %20 into + and decoding in Form mode
turns + back into a space. Only use it for query strings and form bodies. It swaps spaces and does
nothing else: the browser’s URLSearchParams serialiser also escapes ! ' ( ) ~, which this method
leaves as they are.
Python draws the same line: urllib.parse.quote("a b/c", safe="") gives a%20b%2Fc, and
quote_plus("a b/c") gives a+b%2Fc. Note that quote keeps / by default, so pass safe="" when
encoding a value.
Double encoding: %2520 and friends
Encoding text that is already encoded turns every % into %25:
encodeURIComponent("a b"); // "a%20b"
encodeURIComponent(encodeURIComponent("a b")); // "a%2520b"
The giveaway is %25 followed by two hex digits. It usually comes from encoding at two layers, for
example a client library that encodes parameters and a hand-written encodeURIComponent around the same
value, or a redirect URL that is encoded once as a parameter and again when the redirect is built.
Decoding once returns a%20b, which is why the page shows percent signs instead of spaces. The fix is
to find the second layer and remove it, not to decode twice at the end. The tool warns when the input to
Encode already contains %XX escapes, and when a decoded result still contains them.
“URI malformed” has two different causes
decodeURIComponent throws URIError: URI malformed for either of these, and the message does not say
which:
- A
%that is not followed by two hex digits.decodeURIComponent("100%")throws. This is text that was never encoded, such as a discount code, pasted into a decoder. Escape it as%25first. - Escapes that are valid hex but not valid UTF-8.
decodeURIComponent("%C4%E3%BA%C3")throws, because those are the GBK bytes of 你好. JavaScript only decodes UTF-8, so a link produced by an older system or another character set can never be decoded with it.
The tool reports the position of the problem and which case it is. For the second case, turn on Keep
invalid sequences to decode the valid parts and leave the rest as they are, and then work out the source
charset. Encoding can fail too: a lone half of an emoji (an unpaired surrogate such as "\uD800") makes
encodeURIComponent throw, because it cannot become UTF-8.
Chinese and emoji are bytes, not characters
Percent-encoding works on bytes, so what a character becomes depends on the charset that turns it into
bytes. In UTF-8, 你好 is six bytes, %E4%BD%A0%E5%A5%BD, and an emoji such as 😀 is four,
%F0%9F%98%80. The same word in GBK is four bytes, %C4%E3%BA%C3. Both are legitimate percent-encoding
and neither is wrong, but the receiver has to use the charset the sender used. Modern browsers and the
URL standard use UTF-8, so a server that decodes with a legacy default will show garbage for anything
outside ASCII; the fix belongs on the server’s URL or request charset setting.
Avoid escape(). It is deprecated and not UTF-8: escape("中") returns %u4E2D, a form that no standard
URL parser accepts.
Using the URL encoder
Pick Encode or Decode, then the method: Component for a single value, Whole URL to keep : / ? # & =,
or Form for + spaces. The status line counts the escapes it converted, and the Parts of the link
table splits a full address into scheme, host, path and a decoded table of query parameters. That table
is a quick way to confirm that a value containing & stayed in one parameter. It decodes each value the
way a form parser would, so a + in the input shows as a space there. Conversion runs in the page with the
browser’s own functions, and nothing is sent anywhere.