All articles

URI vs URL: The Difference Explained with Examples

Understand the difference between URI, URL and URN: RFC 3986 definitions, generic syntax, schemes, relative references and real code examples.

Neon web address branching into identifier, locator and name components

Introduction

Ask ten developers to explain the difference between a URI and a URL and you will probably get ten slightly different answers. Some say the terms are interchangeable. Some say a URL always starts with http. Some remember a rule about "identify versus locate" but cannot say why anyone bothered to invent two words for something that looks like one string of text.

The confusion has a historical cause. In 1994, RFC 1630 described a Universal Resource Identifier and already predicted the split: some identifiers would carry enough information to find a resource, and others would only name it. That same year RFC 1738 defined Uniform Resource Locators, and RFC 1737 set out requirements for Uniform Resource Names. Three related ideas were born within months of each other, and the vocabulary has been leaking ever since.

By January 2005, RFC 3986 settled the architecture: URI is the general concept, and "URL" and "URN" are informal labels for what a particular URI does. That single sentence resolves most arguments — but only if you understand the syntax underneath it, which is exactly what this article unpacks.

This matters far beyond trivia. The distinction shows up whenever you configure an OAuth redirect_uri, declare an XML namespace, write a Java URI class instead of URL, define a REST resource identifier, or debug why a relative link resolved to the wrong path. Read on and the next time a code review argues about naming, you will be the one holding the specification.

One generic syntax, two familiar roles

A URI is a compact sequence of characters that identifies an abstract or physical resource. The resource can be a web page, but it can equally be a book, a phone number, a person, a namespace or a concept that no network can deliver. Identification is the only requirement.

A URL is a URI that additionally tells you how to get the resource, by describing its primary access mechanism — typically a network protocol and location. Because every URL identifies something, every URL is a URI. The reverse is not true.

A URN is a URI intended to remain a globally unique, persistent name even if the resource moves or disappears. RFC 8141 defines the urn scheme with the form urn:namespace:specific-string, as in urn:isbn:0451450523. It names the book; it does not tell you which server holds a copy.

RFC 3986 deliberately stops using "URL" and "URN" as formal categories. It treats them as descriptions of usefulness rather than distinct types, because a single string can behave as a locator today and remain a stable identifier for decades. The specification defines only URI, its generic syntax and its resolution rules.

Technical foundations and behavior

Everything in the URI family shares one grammar, defined in RFC 3986 as scheme, authority, path, query and fragment. Understanding which components are present, and which are optional, explains almost every practical difference you will encounter between a URI, a URL and a relative reference.

The syntax is: scheme ":" hier-part [ "?" query ] [ "#" fragment ]. The scheme is mandatory in an absolute URI and is the part that determines how everything after the colon must be interpreted.

Every absolute URI needs a scheme

The scheme is the leading token before the first colon: http, https, ftp, mailto, tel, file, urn, data, git, jdbc. It is case-insensitive but conventionally lowercase. Without a scheme you do not have a URI at all — you have a relative reference, which only becomes a full URI after being resolved against a base.

The authority is what makes a URI look like a URL

When a scheme uses a network location, the hierarchical part begins with a double slash and carries an authority: optional userinfo, a host and an optional port. That authority component is precisely the "where to find it" information that earns a URI the informal label URL. mailto:team@utily.tools has no authority and is not a URL; https://utily.tools/url-codec has one and is.

Path, query and fragment carry the rest of the meaning

The path is hierarchical data identifying the resource within the authority. The query, introduced by ?, carries non-hierarchical data such as filters or parameters. The fragment, introduced by #, identifies a secondary resource within the primary one and is never sent to the server — the browser resolves it locally, which is why analytics tools cannot see it.

URI reference: absolute, relative and same-document

RFC 3986 defines a URI reference as either a URI or a relative reference. /articles/uri-vs-url-difference-explained, ../assets/logo.png and #section-3 are all valid references but none is a URI on its own. Section 5 of the specification gives the exact algorithm for merging them with a base URI — the algorithm every browser and HTTP client implements.

Reserved characters make the distinction operational

Because the delimiters :, /, ?, #, [, ], @, &, = and + carry structural meaning, any user data placed inside a component must be percent-encoded first. This is where the theory becomes a bug report: an unencoded ampersand inside a query value silently becomes a new parameter. Verify such values in the utily.tools URL Codec before shipping them.

IRI, and why non-ASCII addresses still work

RFC 3986 restricts URIs to a subset of ASCII. RFC 3987 defines the Internationalized Resource Identifier, which permits Unicode and maps back to a URI by encoding path and query bytes as UTF-8 percent triplets and converting host labels through IDNA Punycode. What you see in the address bar is often an IRI; what travels on the wire is a URI.

Real-world applications

OAuth 2.0 redirect_uri

The specification names the parameter redirect_uri, not redirect_url, and requires an absolute URI. Native apps legitimately register custom schemes such as com.example.app:/callback, which identify a handler rather than a network location.

XML and RDF namespaces

A namespace declared as http://www.w3.org/1999/xhtml is a URI used purely as a unique name. Dereferencing it is optional and irrelevant to parsing — a perfect example of a URI that looks like a URL but is not used as one.

Java URI versus URL

java.net.URI parses and normalizes strings without network assumptions, while java.net.URL can open a connection and, in older versions, performs DNS lookups inside equals(). Choosing the right class prevents surprising latency and comparison bugs.

REST resource identifiers

An API contract that says "the order is identified by /orders/{id}" is describing a URI reference; the deployed https://api.example.com/orders/42 is the URL that resolves it. Keeping the two apart makes versioning and environment promotion far cleaner.

Persistent identifiers

A DOI expressed as doi:10.1000/182 or an ISBN as urn:isbn:0451450523 stays valid after any migration, while the URL that serves the file may change several times over the same period.

Standards and deeper technical reference

Where the words came from: 1994 and the birth of three acronyms

When Tim Berners-Lee documented the addressing scheme of the early Web in RFC 1630 in June 1994, the document was titled "Universal Resource Identifiers in WWW". It already made the key observation that would define the next thirty years of debate: some identifiers encode enough information to locate a resource, while others merely name it. The paper described the former as locators and the latter as names.

Later that year, RFC 1738 formalized Uniform Resource Locators, listing the schemes that browsers actually used — ftp, http, gopher, mailto, news, telnet, file. RFC 1737 published functional requirements for Uniform Resource Names in December 1994, describing an identifier that should be globally unique, persistent and independent of location. The vocabulary of URI, URL and URN was therefore complete before most people had ever used a browser.

RFC 2396 replaced "Universal" with "Uniform" in August 1998 and unified the grammar. Its successor, RFC 3986, arrived in January 2005 as STD 66 and made a decision that most casual explanations still miss: it stopped treating URL and URN as formal classes. Section 1.1.3 explains that the terms are useful descriptions but that the specification defines only URI, since a single identifier can be both a durable name and a working locator at the same time.

Milestones in URI standardization
DateDocumentContribution
June 1994RFC 1630First description of Universal Resource Identifiers and the locator/name split
December 1994RFC 1738Defined Uniform Resource Locators and the original scheme list
December 1994RFC 1737Functional requirements for Uniform Resource Names
May 1997RFC 2141Original URN syntax
August 1998RFC 2396Unified generic URI syntax; "Uniform" replaces "Universal"
January 2005RFC 3986 (STD 66)Current generic syntax; URL and URN become informal descriptions
January 2005RFC 3987Internationalized Resource Identifiers (IRIs) with Unicode
April 2017RFC 8141Current URN syntax, replacing RFC 2141
Living standardWHATWG URLBrowser-accurate parsing model for web URLs

The generic syntax, component by component

RFC 3986 defines a URI as scheme ":" hier-part [ "?" query ] [ "#" fragment ]. That single production covers every identifier you will ever meet, from https://utily.tools/url-codec?value=a%20b#result to urn:isbn:0451450523 and tel:+551140028922. What changes between them is which optional components appear, not the grammar itself.

Consider https://user:pass@api.utily.tools:8443/v2/orders/42?fields=id,total#summary. The scheme is https. The authority is user:pass@api.utily.tools:8443, itself split into userinfo, host and port. The path is /v2/orders/42. The query is fields=id,total. The fragment is summary. Remove the authority and you have something like mailto:team@utily.tools — still a perfectly valid URI, just not a URL.

A crucial and frequently misunderstood detail: the fragment is a client-side concern. RFC 3986 section 3.5 states that it identifies a secondary resource by reference to the primary resource, and the dereferencing is performed by the user agent. HTTP requests never carry it. This is why server logs cannot show which anchor a visitor jumped to, and why single-page applications historically used fragments for routing without any server round trip.

Components of https://user:pass@api.utily.tools:8443/v2/orders/42?fields=id,total#summary
ComponentValueDelimiterRequired?
Schemehttpsfollowed by :Yes, in an absolute URI
Userinfouser:passfollowed by @No, and deprecated for credentials
Hostapi.utily.toolsafter //Only when an authority exists
Port8443preceded by :No; scheme default applies
Path/v2/orders/42starts with /Yes, but may be empty
Queryfields=id,totalpreceded by ?No
Fragmentsummarypreceded by #No; never sent to the server

Identifying versus locating: the test that actually works

The reliable test is not whether the string starts with http. It is whether the identifier describes an access mechanism. If it tells an agent which protocol to speak and which endpoint to speak it to, it functions as a locator. If it only asserts uniqueness, it functions as a name.

This is why some examples surprise people. The XML namespace http://www.w3.org/1999/xhtml uses the http scheme and syntactically qualifies as a URL, yet it is used exclusively as a name — no conforming parser fetches it, and the W3C has published guidance explaining that namespace URIs need not be dereferenceable. Conversely urn:isbn:0451450523 is unambiguously a name, and resolving it to a copy of the book requires an external resolution service.

The pragmatic conclusion drawn by RFC 3986 is that classification depends on use, not on shape. A given string can serve as a durable identifier in a contract and as a working address in production, which is exactly why the standard defines only one term and lets context supply the rest.

URI
Identifies a resource. The umbrella concept; the only one formally defined by RFC 3986.
URL
A URI that also specifies a primary access mechanism, normally a network protocol plus location.
URN
A URI under the urn scheme, intended to be a globally unique and persistent name.
URI reference
Either a full URI or a relative reference that must be resolved against a base.
IRI
An internationalized identifier permitting Unicode, mapped to a URI for transmission.

Relative references and the resolution algorithm

Most links inside a real application are not URIs at all. /articles/uri-vs-url-difference-explained, ../assets/cover.png, ./next and the bare #summary are relative references. RFC 3986 section 5 specifies precisely how a base URI and a reference combine, and every browser, HTTP client and crawler implements the same steps.

The algorithm resolves in order: if the reference has a scheme it is already absolute; otherwise inherit the base scheme. If it has an authority, keep it and take the path as given. If the path starts with a slash, replace the base path entirely. If the path is empty, keep the base path and substitute the query when one is present. Otherwise merge the reference onto the base path directory and run remove_dot_segments to collapse . and .. sequences.

Two consequences trip teams up constantly. First, a base of https://utily.tools/articles/guide with reference "next" resolves to /articles/next, not /articles/guide/next — the last segment is treated as a file, not a directory, unless the base ends in a slash. Second, a reference beginning with a double slash, such as //cdn.example.com/lib.js, is a protocol-relative reference that inherits only the scheme and replaces the entire authority. That form was once idiomatic and is now discouraged, since it silently follows an insecure base into HTTP.

Resolution against base https://utily.tools/articles/guide?x=1
ReferenceResolved URIRule applied
otherhttps://utily.tools/articles/otherMerge with base directory
/url-codechttps://utily.tools/url-codecAbsolute path replaces base path
../abouthttps://utily.tools/aboutMerge then remove dot segments
?y=2https://utily.tools/articles/guide?y=2Empty path keeps base path, query replaced
#tophttps://utily.tools/articles/guide?x=1#topSame-document reference
//cdn.example.com/a.jshttps://cdn.example.com/a.jsScheme inherited, authority replaced
mailto:hi@utily.toolsmailto:hi@utily.toolsReference already has a scheme

Normalization, equivalence and why two URIs can be the same

Comparing identifiers naively causes cache misses, duplicate database rows and broken authorization checks. RFC 3986 section 6 defines a ladder of comparison strategies, from simple string equality to scheme-based normalization, and warns that false negatives are usually more damaging than the cost of normalizing.

Syntax-based normalization is always safe: lowercase the scheme and host because they are case-insensitive, uppercase the hexadecimal digits in percent triplets, decode percent-encoded octets that correspond to unreserved characters, and remove dot segments from the path. Scheme-based normalization adds knowledge of defaults — https://utily.tools:443/ equals https://utily.tools/, and an empty path in an http URI is equivalent to /.

What is never safe to assume is that the path is case-insensitive, that a trailing slash is meaningless, or that reordering query parameters preserves identity. All three depend entirely on the server. This is the layer where SEO and engineering meet: canonical link elements exist precisely because a site can serve identical content under several non-equivalent URIs, and search engines need to be told which one is authoritative.

Case normalization
Scheme and host lowercase; percent-encoding hex digits uppercase.
Percent-encoding normalization
Decode triplets representing unreserved characters such as %7E to ~.
Path segment normalization
Apply remove_dot_segments to eliminate . and .. sequences.
Scheme-based normalization
Drop default ports and treat an empty http path as /.
Never assume
Path case-insensitivity, trailing-slash equivalence or query parameter order.

The URN scheme in detail

RFC 8141 defines the current URN syntax as urn:NID:NSS, optionally followed by r-components, q-components and an f-component. The namespace identifier is registered with IANA; the namespace-specific string is governed by whoever owns that namespace. Case matters: the leading urn token and the NID are case-insensitive, but the NSS generally is not.

Real namespaces include isbn for books, uuid for identifiers such as urn:uuid:6e8bc430-9c3a-11d9-9669-0800200c9a66, ietf for specifications, and oasis for OASIS standards. The persistence guarantee is a social and organizational commitment, not a technical property — a URN survives migrations only because a registry keeps its meaning stable.

This is why URNs are common in publishing, library science, telecommunications and legal citation, and rare in day-to-day web development. When a system does need permanence, the usual pragmatic compromise is a persistent HTTP URI backed by a redirect service, as with DOIs served through doi.org. That approach gains dereferenceability at the cost of depending on one organization keeping the resolver alive.

Comparing the three roles side by side
PropertyURI (general)URLURN
PurposeIdentify a resourceIdentify and locateIdentify with a persistent name
Defined byRFC 3986RFC 1738, described by RFC 3986RFC 8141
Access mechanismNot requiredRequiredNot provided
Examplemailto:team@utily.toolshttps://utily.tools/url-codecurn:isbn:0451450523
Survives relocationDependsOften breaksBy design
Directly fetchableDependsYesOnly via a resolver

RFC 3986 versus the WHATWG URL Standard

There are two living specifications, and they do not agree. RFC 3986 gives an ABNF grammar for identifiers in general. The WHATWG URL Standard describes a state-machine parser that reproduces what browsers actually do with web URLs, including error recovery for input that RFC 3986 would simply reject.

The divergences are concrete. WHATWG parsing has no notion of an empty-but-present authority in the RFC sense, applies IDNA and percent-encode sets per component, resolves backslashes as slashes in special schemes, and defines an "opaque path" for non-special schemes. Given http://example.com/a\b, a browser normalizes the backslash to a forward slash; an RFC 3986 parser treats it as an invalid path character.

The engineering rule that follows is straightforward. Use the WHATWG model when the identifier will be handled by a browser, and the RFC 3986 model for protocol identifiers, namespaces and general-purpose URIs. Never rebuild an address by concatenating strings across the two models, and never let one component pass through a parser configured for another — this is the root cause of a large family of parser-confusion vulnerabilities.

Security: where the distinction becomes an incident report

URI parsing differences are a documented attack surface. When a proxy, a framework and an application each parse the same string with slightly different rules, an attacker can craft input that one layer treats as safe and another treats as a different resource entirely. Research on URL parser confusion has repeatedly shown request smuggling and access-control bypasses arising from exactly this mismatch.

Server-side request forgery depends on the same ambiguity. A validator that checks whether a URL "contains" an allowed host is trivially defeated by https://allowed.example.com@attacker.example/ — the userinfo component makes attacker.example the real authority. Correct validation parses the URI, extracts the host component structurally and compares it against an allowlist, never using substring matching.

Open redirects follow from the protocol-relative form: a redirect target of //attacker.example passes a naive "must start with /" check while sending the user off-site. Encoded traversal sequences such as %2e%2e%2f exploit decode-order mistakes when a path maps to a filesystem. In every case the mitigation is the same discipline: parse structurally, normalize exactly once at a defined boundary, validate the canonical form, and reject anything malformed rather than repairing it.

Parse, do not pattern-match
Extract host, path and query with a real parser before validating anything.
Beware userinfo
https://trusted.example@evil.example/ has authority evil.example.
Reject protocol-relative redirects
//host bypasses checks that only require a leading slash.
Normalize once
Decoding twice turns %252e%252e into .. after the security check has passed.
Encode per component
Path, query and fragment have different permitted character sets.

Practical guidance for APIs, code and documentation

In API contracts, prefer "URI" when you mean identity and "URL" when you promise retrieval. An OpenAPI description that says an order is identified by /orders/{id} is describing a URI reference that stays valid across every environment; the deployment-specific https://api.example.com/orders/42 is the URL. Mixing them bakes hostnames into contracts and makes environment promotion painful.

In code, choose the type that matches the intent. Java draws the line sharply: java.net.URI parses and normalizes without network semantics, while java.net.URL can open connections and, in older releases, performed DNS resolution inside equals() and hashCode() — a well-known source of unexpected latency. In JavaScript, the URL constructor implements the WHATWG model and throws on invalid input, which makes it a reasonable validation gate. In Python, urllib.parse.urlsplit follows RFC 3986 closely and preserves components rather than guessing.

For content and SEO, the practical items are ordinary but effective: pick one canonical form and declare it with a rel=canonical link, keep paths readable and lowercase, avoid duplicate content served under both trailing-slash variants, and remember that fragments are never seen by a crawler as separate pages. None of this requires arguing about terminology — it just requires knowing which component you are changing.

When you need to inspect what a component actually contains, decode it rather than guessing. The utily.tools URL Codec percent-encodes and decodes values directly in the browser, which makes it quick to confirm whether a redirect_uri was double-encoded, whether a query value swallowed an ampersand, or how a UTF-8 string is represented on the wire.

Frequently confused cases, answered directly

Is every URL a URI? Yes. Is every URI a URL? No — mailto:team@utily.tools, urn:isbn:0451450523 and tel:+551140028922 identify resources without describing a retrievable network location. Does a URI have to start with http? No; the scheme is whatever the identifier declares, and thousands are registered with IANA.

Is /articles/guide a URL? Strictly, no. It is a relative reference, and it only becomes a URI after resolution against a base. Is a URN a URL? No, though a resolver can map one to a URL. Is a file path a URI? Only when expressed with the file scheme, as in file:///home/user/notes.txt.

Why does OAuth say redirect_uri instead of redirect_url? Because the value need not be an HTTP address — mobile applications legitimately register custom schemes such as com.example.app:/callback. And why does the WHATWG standard barely mention "URI"? Because it models what browsers do with web addresses specifically, and chose the more familiar term for that narrower job.

Primary specifications and references

Conclusion

The hierarchy is simple once stated precisely: URI is the general concept of identifying a resource, URL is the informal name for a URI that also describes how to access it, and URN is the informal name for a URI meant to be a persistent name. Every URL is a URI; not every URI is a URL. Since RFC 3986, the specification itself only defines URI and its generic syntax of scheme, authority, path, query and fragment.

In my view, the practical value of the distinction is not vocabulary policing but design clarity. When you say "URI" you are committing only to identity, which keeps contracts, namespaces and API documentation independent of hosting decisions. When you say "URL" you are promising a retrieval path, and that promise has an expiry date. Teams that blur the two end up hard-coding environments into identifiers.

A closing thought worth carrying forward: the strings you write in configuration files are not just text — they are a small grammar with reserved characters that can change a request structure the moment user data slips in unencoded. Try encoding and decoding a few real values in the utily.tools URL Codec, then continue with the article on URL percent-encoding to see exactly where that grammar breaks and how to protect it.

Open URL Codec Read more articles