What does a URL mean? The string of characters in the address bar is broken down to five segments, and the content after the hashtag simply doesn't leave the web

last week, my sister sent me a string of URLs on WeChat, saying she absolutely couldn't open them and asked me to take a look. I laughed when I saw it: the URL was half-broken when she copied it, the path was down to just the first letter, and the rest was gone. She broke the WeChat URL in the middle when she copied it. I sent the full URL back, and she opened it immediately. She even asked me: Which paragraph in this string is important, and which is dispensable?
this question is really worth discussing. The string of things in the address bar is technically called URL. I opened up my blog's URL and tested it overnight. Each of the five sections managed their own tasks, and after figuring it out, when I encountered "URLs won't open," I could basically tell at a glance where the link was disconnected.
first dissected a string of URLs: five sections each have their own section
using an article from my own blog as a sample:
https://www.shunot.com/lybk/1069.html
Split from left to right: https:// is the protocol, www.shunot.com is the domain name, and a port (443) is omitted. /lybk/1069.html is the path. If the URL were a bit longer, you could add ? parameter and #anchor point , for example, https://www.shunot.com/lybk/1069.html?v=2#comment-8 and you'd get the complete five-paragraph format.
to put it another way: the protocol is whether you drive or walk when you go out, the domain name is the community address, the port is the building number, the path is the number of rooms, the question mark is the message you sent to the owner, and the donkey is your own memorization of "Go straight to the fridge upon entering." The owner can see the message, but the idea of rushing to the fridge is something the owner doesn't know at all—this metaphor will come true later in real life.
| a paragraph looks like | what does it matter | what happens if you type it wrong | |
| https: | use encrypted lanes or naked lanes | intranet devices often fail to open | |
| domain | www.shunot. com | find a server | you can directly jump to the search page |
| port | :443 or :8080 | which server door to go through | connection timeout spinning |
| path | /lybk/1069.html | Which file | 404 page |
| question mark/hash | ?v=2#comment-8 | message/self-mark | most likely still open |
protocol, domain name, and port: The first half of the tube finds the server
starts with an unconventional test. I changed all domain names to uppercase HTTPS://WWW. SHUNOT.COM/lybk/1069.html Request, return 200, exactly the same as before. Because domain names don't matter if it's case-sensitive; you can type it Www.ShUnOt.CoM fine. domain name is just for reporting the address to the server, doesn't matter if you report it.
Protocol Header usually doesn't need to be thrown by hand. You type shunot.com and press Enter in the address bar, and the browser automatically adds https:// and www.—that's what it does. But this automatic patching on internal network devices causes problems: when accessing the old router's backend and typing 192.168.1.1, the new browser will default to the https://, while the old device doesn't support encryption at all and just spins a white screen. At this point, you need to manually type all four letters http://. I wrote about this pitfall in the about the can't access the backend, especially on phones—typing 192.168.1.1 brings up a search results page. That's the input method using the address as a search keyword. The solution is in the of the phone login backend article.
https Compared to HTTP, it has an extra lock and encrypts the transmission process. In which situations should you be concerned about this lock, and which don't matter? I compared the packet capture before, and the plaintext clearly shows even the server version number. article explains it in more detail. Between
domain name and port, there's a hidden segment: port. You might never have seen it on the website because it usually doesn't show up—https defaults to 443, http defaults to 80, and browsers save you the trouble. When must a port be written? When the server is not on the default door.
my NAS music server runs on port 8080, you have to write http://192.168.31.20:8080 to access it, and skip the part after the colon. If you go to the browser and knock on door 80, no one answers. I've written a dedicated article about the specifics of this "house number" set: who sets the numbers 3389, 5005, 554, and why the external port and internal port are two boxes when mapping ports—all covered there. The symptoms of the
port error are quite recognizable: it's not 404, it's just a constant loop and finally times out. Because 404 is at least "the door is open but the file is gone," and the wrong port means "this door isn't open at all," and the server won't even respond to you.
the path section: a single mistake in a letter changes the character, case sensitivity is
path is the most meticulous of the five segments. I tested changing /lybk/ to /LYBK/ to request it, and the server returned a 301 jump, returning a lowercase address to me; Switching to /Lybk/ also makes it 301. My blog's server is pretty good-natured; after implementing compatibility redirects, with a few more detours, you can still get there. But most websites' servers aren't so generous—web servers running Linux have case-sensitive file systems, and the path is two different files, so you can just throw a 404 right in your face.
my sister copied the broken URL string, the bad part is the path section: the domain name remains, but the path only has the beginning, so naturally nothing can be found. Therefore a website can't be opened, first check whether the path is complete , especially long URLs copied from WeChat or SMS, which are most likely to be broken or cut off.
By the way, domain names and IPs can also be used interchangeably. Domain names are essentially translation systems invented to help people remember IPs. You can access them directly by entering the IP address, but nowadays, most websites have multiple sites with one IP address, separated by domain names, so direct IP access often ends up on the default page. IP and domain names have their own secrets I've written about it before, and the article explaining why external networks can't connect to those starting with 192.168 is thorough.
Question marks and hash points: one sent to the server, the other only on your computer
These are the two most interesting brothers in the five segments—they look alike, but their fates are completely different.
Query parameters are called after the question mark. I tested it by adding a ?v=20261004 to the end of the URL, and when I requested 200, the content was exactly the same as before—because my server didn't care about this argument at all and gave me the page as is. But the parameter is not just for show: the search page's ?q = keyword, the video site's ?v = video number—all rely on this. You send the search result URL to someone else, and the segment after the question mark is their search term, so send it all over together.
the password after the hashtag is called the anchor point. This is a real trivia: I glanced at the packet capture and saw the URL with #comment-8 to request it. The server received the request line GET /lybk/1069.html — the content after the hashtag was never online at all, it only exists in your own browser, used to jump to a certain location after the page loads. Direct links to comment sections and page directory redirects are all done by it.
this feature has a practical trick: when the webpage gets stuck and old content won't refresh, you can casually add a ?123 to the end of the URL, then press Enter, which tricks the browser into thinking "this is a new URL" and bypasses the cache to reload it. Of course, the proper method is Ctrl+F5, how to clear four-layer cache. I wrote an article the cache pit in WeChat's built-in browser is also in that article.
the password string with the percent sign: space becomes %20, the password with # is a pitfall. I went through it for an hour
The website cannot contain spaces or Chinese characters, these "special characters"; they will be translated into the three-digit cipher text starting with the percent sign: space is %20, and one Chinese character can become nine characters. I tested that I put a bare space in the path to request it, and it couldn't connect at all; After switching to %20 the request is sent normally, returning 404—this means the server has received it and restored %20 to a space to search for files, except the file with the space doesn't exist.
the most frustrating scenario of this escape is the cipher. I assigned the traffic link to my cousin's surveillance camera, and the password had a call #. No matter how I entered it, it wouldn't connect. After nearly an hour of trouble, I finally realized that the hash in the URL was a 'start of anchor' mark, and the server only read the password before the hash and cut it off. Replace # with %23 and fill in again, instant connection. In cases where passwords are embedded in URLs, such as camera stream addresses or NAS direct links, the hashtags, IDs, and spaces in the password must be interpreted according to this set of rules.
On the other hand, if you see phrases like %20, %23, %E4%BB%80 in the address bar, don't panic—those are the website address disguises with spaces, hash numbers, and Chinese characters, which the browser automatically adds when you copy and paste them.
Ending: If the website can't be opened, scan five sections in this order
| symptoms | first check which section | most likely cause |
| a search page | protocol/domain name pops up | address is treated as a search term |
| white screen spins in circles to timeout | port | port is either not written or written incorrectly |
| 404 page | path | copying half-broken or case- |
| pages is an old | question mark segment | cache, add ?123 or Ctrl+F5 |
| to open but don't jump to the specified position | Lot number segment | anchor missing or page redesign |
my sister now knows how to paragraph URLs. Last month, she even found the word 'utm_source' in a website and asked if it tracked where she came from—those UTM-prefix parameters after the question mark, most likely for this. Delete them and visit again, the page opens as usual, and the world is quiet.
