encodeURIComponent vs encodeURI, and where %2520 and "URI malformed" come from

Updated 5 min read

Percent-encoding looks like one operation, but there are three common variants, and mixing them up produces the same handful of bugs: a query that loses half its parameters, a %2520 in the address bar, a URIError, and Chinese text turning into question marks. Each has one cause.

Which characters each function leaves alone

encodeURIComponent leaves only letters, digits and - _ . ! ~ * ' ( ). encodeURI leaves those plus the characters that structure an address: ; , / ? : @ & = + $ #.

CharacterencodeURIComponentencodeURI
/%2Fkept
?%3Fkept
#%23kept
&%26kept
=%3Dkept
+%2Bkept
%%25%25
space%20%20
中%E4%B8%AD%E4%B8%AD
encodeURIComponent("a b&c=中"); // "a%20b%26c%3D%E4%B8%AD"
encodeURI("a b&c=中");          // "a%20b&c=%E4%B8%AD"

The rule that follows: encodeURIComponent is for one piece that will sit inside a URL, such as a query value or a path segment. encodeURI is for an address that is already assembled, where you only want spaces and non-ASCII characters escaped and the separators left working.

A value with & or = splits the query

The symptom is a parameter that arrives truncated, or an extra parameter nobody sent. It appears when a value is built with encodeURI, or with no encoding at all:

const q = "R&D = 50% off";
"/search?q=" + encodeURI(q);          // /search?q=R&D%20=%2050%25%20off
"/search?q=" + encodeURIComponent(q); // /search?q=R%26D%20%3D%2050%25%20off

The first form is read as q=R plus a second parameter named D . The second keeps the whole phrase in q. encodeURI cannot help here because it has no way to know that the & is data rather than a separator. Encode each value separately, then join with literal & and =. Or skip the manual work: new URLSearchParams({ q }).toString() does the escaping for you.

The same applies to #. An unescaped # in a value starts the fragment, and everything after it is never sent to the server.

Space is %20 in one place and + in another

%20 is correct everywhere in a URL. + means a space only inside a query string or a application/x-www-form-urlencoded body, which is what an HTML form submits. In a path, + is a literal plus sign. Decoding shows the difference:

decodeURIComponent("a+b");                  // "a+b"
new URLSearchParams("a=1+2").get("a");      // "1 2"

Two bugs come from this. A value containing a real + (a phone number, c++, a base64 string) loses it when it is encoded with a form-style encoder that does not escape it, because the receiver turns it into a space; it has to travel as %2B. And a + that ends up in a path stays a plus, so a “space” you meant there is wrong. In the URL encoder, the Form method turns %20 into + and decoding in Form mode turns + back into a space. Only use it for query strings and form bodies. It swaps spaces and does nothing else: the browser’s URLSearchParams serialiser also escapes ! ' ( ) ~, which this method leaves as they are.

Python draws the same line: urllib.parse.quote("a b/c", safe="") gives a%20b%2Fc, and quote_plus("a b/c") gives a+b%2Fc. Note that quote keeps / by default, so pass safe="" when encoding a value.

Double encoding: %2520 and friends

Encoding text that is already encoded turns every % into %25:

encodeURIComponent("a b");                       // "a%20b"
encodeURIComponent(encodeURIComponent("a b"));   // "a%2520b"

The giveaway is %25 followed by two hex digits. It usually comes from encoding at two layers, for example a client library that encodes parameters and a hand-written encodeURIComponent around the same value, or a redirect URL that is encoded once as a parameter and again when the redirect is built. Decoding once returns a%20b, which is why the page shows percent signs instead of spaces. The fix is to find the second layer and remove it, not to decode twice at the end. The tool warns when the input to Encode already contains %XX escapes, and when a decoded result still contains them.

“URI malformed” has two different causes

decodeURIComponent throws URIError: URI malformed for either of these, and the message does not say which:

  • A % that is not followed by two hex digits. decodeURIComponent("100%") throws. This is text that was never encoded, such as a discount code, pasted into a decoder. Escape it as %25 first.
  • Escapes that are valid hex but not valid UTF-8. decodeURIComponent("%C4%E3%BA%C3") throws, because those are the GBK bytes of 你好. JavaScript only decodes UTF-8, so a link produced by an older system or another character set can never be decoded with it.

The tool reports the position of the problem and which case it is. For the second case, turn on Keep invalid sequences to decode the valid parts and leave the rest as they are, and then work out the source charset. Encoding can fail too: a lone half of an emoji (an unpaired surrogate such as "\uD800") makes encodeURIComponent throw, because it cannot become UTF-8.

Chinese and emoji are bytes, not characters

Percent-encoding works on bytes, so what a character becomes depends on the charset that turns it into bytes. In UTF-8, 你好 is six bytes, %E4%BD%A0%E5%A5%BD, and an emoji such as 😀 is four, %F0%9F%98%80. The same word in GBK is four bytes, %C4%E3%BA%C3. Both are legitimate percent-encoding and neither is wrong, but the receiver has to use the charset the sender used. Modern browsers and the URL standard use UTF-8, so a server that decodes with a legacy default will show garbage for anything outside ASCII; the fix belongs on the server’s URL or request charset setting.

Avoid escape(). It is deprecated and not UTF-8: escape("中") returns %u4E2D, a form that no standard URL parser accepts.

Using the URL encoder

Pick Encode or Decode, then the method: Component for a single value, Whole URL to keep : / ? # & =, or Form for + spaces. The status line counts the escapes it converted, and the Parts of the link table splits a full address into scheme, host, path and a decoded table of query parameters. That table is a quick way to confirm that a value containing & stayed in one parameter. It decodes each value the way a form parser would, so a + in the input shows as a space there. Conversion runs in the page with the browser’s own functions, and nothing is sent anywhere.

Open the tool: URL Encoder & Decoder

Back to guides

More guides