This article, and most articles about this, doesn't explain where FBI got that GDID from. Ok, Microsoft has a list of IP addresses that has been used by a computer with a certain GDID, but FBI needs to get the GDID in the first place, and then try to bind that to a person.
I found another article that explains the process a bit better:
> Stokes got caught because he used the same Windows device for everything, and the GDID stitched all of it back together after the fact.
> Scattered Spider members phoned the jewelry retailer’s IT help desk from Google Voice numbers, posed as locked out employees, and talked support staff into resetting three accounts, two with administrator privileges. From there they installed a tunneling tool called ngrok to get past the retailer’s network defenses, moved roughly 77 gigabytes of data to Amazon cloud storage using ngrok [...]
> Investigators later subpoenaed ngrok and found the account used in the attack had been created on May 12, 2025, at 19:21 UTC from a VPN proxy IP address run by Tzulo, a hosting provider. The IP was a dead end. VPN proxies do that. But the GDID is built different.
> Microsoft’s records showed that at that exact same minute, a Windows device carrying GDID g:6755467234350028 had visited the ngrok signup page. Three hours later, the same GDID visited the retailer’s own website, through the same Tzulo proxy address used to set up the ngrok account. It gave the FBI a device, that don’t rotate the way VPN exit nodes do.
https://www.windowslatest.com/2026/07/10/you-cant-fully-disa...
Although this doesn't explain where Microsoft got that traffic data from. How do Microsoft know which sites a computer visit?
What they did was the opposite: ask Microsoft for GDIDs used by attacker-associated IPs within several 24-hour time periods during which attack-related activity took place. Windows pings Microsoft regularly with the GDID, establishing links between your GDID and any IP addresses you use. The IP logs from Microsoft and the VPS provider showed at least 10 instances where a single VPN IP accessed the attacker's VPS and also pinged Microsoft with at least one GDID within a 24-hour period. They found a constant GDID that all instances shared. This seems to have been the most damning GDID-related evidence in the DOJ complaint [1] and yet it wasn't mentioned in the article you linked (or any other articles about this I've seen pop up on HN). It includes the diagram from the complaint (page 18) that outlines this, but devoid of context. The ngrok stuff that the article focuses on was just the cherry on top and was discussed later in the complaint.
What also becomes clear when you read the complaint is that the GDID was just one piece of the puzzle and that they had plenty of other evidence. Attacker-associated IPs were used to access the suspect's Apple, Snapchat, and Facebook accounts, at least one of which was his actual residential IP, not a VPN IP. Once they had revealed the identity of the person who owned these accounts, they were able to all-but-confirm that this was in fact the attacker.
What remains unclear even after reading the complaint is how they were so sure that the GDID they obtained visited specific websites, but honestly, at that point, they were already drowning in evidence, so I don't know if it matters that much. It could be as simple as "he was signed into Edge with his Microsoft account and had sync enabled".
[1] https://www.justice.gov/usao-ndil/media/1450651/dl?inline
* Microsoft collects timestamped GDID - IP address combinations at some interval.
* Other Service (eg your website) collects timestamped IP address data.
Combine the above and you can tell which GDID visited the service without the service knowing anything about GDID.
Eg
At 9:47:00 Microsoft gets a ping from a computer at IP address "123" with a GDID of "abc".
At 9:48:07 mywebsite.com is visited by IP address "123".
You combine both sets of data and you can be reasonably sure that GDID "abc" visited mywebsite.com at 9:48:07.
"required" spyware: https://learn.microsoft.com/en-us/windows/privacy/required-d...
"optional" (but on by default) spyware: https://learn.microsoft.com/en-us/windows/privacy/optional-d...
This is my favorite:
Data Description for Browsing History data type
Microsoft browser data subtype: Information about Address bar and Search box performance on the device
* Text typed in Address bar and Search box
* Service response time
* Autocompleted text, if there was an autocomplete
* Navigation suggestions provided based on local history and favorites
* Browser ID
* URLs (may include search terms)
* Page title and notification textEdit: actually there have been a few threads about this - the link you mentioned was submitted in the second of these:
Microsoft confirms Windows GDID device identifier that cannot be disabled - https://news.ycombinator.com/item?id=48920338 - July 2026 (60 comments)
Microsoft admits Windows 11 has a GDID tracker with no off switch - https://news.ycombinator.com/item?id=48872561 - July 2026 (15 comments)
Windows GDID Changer - https://news.ycombinator.com/item?id=48818707 - July 2026 (4 comments)
Full Writeup of the Windows GDID - https://news.ycombinator.com/item?id=48811081 - July 2026 (49 comments)
Microsoft GDID telemetry includes full browsing and gaming history - https://news.ycombinator.com/item?id=48787239 - July 2026 (6 comments)
> Three hours later, the same GDID visited the retailer’s own website, through the same Tzulo proxy address used to set up the ngrok account.
This criminal mastermind got caught because he did everything but sign his name to the crimes while holding two pieces of government identification in presence of a notary.
The FBI did the bare minimum in terms of old-fashioned detective work, and correlated evidence from various sources.
The obsession with GDID is a complete nothing-burger and I'm tired of seeing it on the front page every other day.
As I grow older, I simply do not have the patience or the will to do all these workarounds and tweaks to my OS to turn it from a piece of barely-working corporate spyware into something that I can call a productive tool.
It also seems like interacting with the OS is on its way out anyway, as most people essentially interact with computers via the browser (essentially a different sandboxed OS altogether), whether on desktop or mobile device.
https://petri.com/windows-10-ignoring-hosts-file-specific-na...
Windows falls back to the normal resolver in that case.
"Oh look, this one has almost all the serial numbers of components and attached-devices as that other one, it's probably the same computer with a fresh install, let's make a note of that..."
It does make me wonder if people would react differently if Linux or Apple did the same thing.
Just stop already, you are in an abusive relationship. Get out. Remove windows. It's not your friend. It's your computer, you have options.
> $lid=(Get-ItemProperty 'HKCU:SOFTWAREMicrosoftIdentityCRLExtendedProperties').LID
which obviously does not work because it has all the slashes removed. Article author apparently didn't bother to proofread anything.
Indeed. One can regenerate it on each shutdown / reboot. The process is a little different depending on whether one has systemd or not. It's a dbus thing but systemd ingests it and there is a specific process around updating that in systemd.
There is also the NetworkID in Firefox about:networking#networkid
Another trackable piece of information on most systems is the creation time of / which just about any application can query unless it is properly isolated. This can be turned into a short unique hash. There are hacky ways to change the Birth time on unmounted filesystems using debugfs which may result in corruption. Some overlay filesystems do not support Birth time but that is not going to help most people unless the general populous expect all applications to be isolated in name spaces and overlay filesystems but this would have to be an expected pattern across all applications universally.
stat / | grep irth
Birth: 2023-04-17 20:27:01.000000000 +0000
stat / | grep irth | md5sum
8b5e849954373f8f3c2a625847e4e858But do any popular linux distributions use an identifier? Ubuntu, Kali, Mint, Arch, etc?
It seems an attractive way for devs to work out telemetry. Awful in reality; but I imagine attractive.
But given there's no cloud accounts on linux I would imagine it's trivially changed just like a NIC's MAC
Also, seems unlikely it would be used for any single-signon with cloud services.
I guess I'm good then.
(Kinda amazed they just didn’t to that instead)
This GDID is essentially a similar concept though, but not as hardware tied, but is reported in telemetry.
when LTSC exists why are people using W11 ?
That's sounds nice and all, but everything about it, from "Nothing is theoretical" to "evidence, tagged by confidence level" goes out the window when you find claude as one of the commit authors, and the content is clearly copy pasted output from claude with very little editing. Worse yet, one of the sources he cites is also clearly AI output.
I'm not even against the use of AI here. I would rather see it clearly say either "this is what claude found after I told it to investigate" or "yes this is generated by claude, but I independently verified each of the points myself".
Microsoft is a post-privacy corporation. Eventually, we really have to stop acting so shocked about this sort of thing from them.
If they'd done their VPN/dirty work in a linux VM/container, with the VPN running in that, they'd have been fine.
Clean-them would have had their GDID on their normal ISP IP or equivalent, and Dirty-them would have had everything through the VPN from their no-telemetry dirty-host.
Ultimately there's no informed consent here: The average consumer is disbelieving and surprised if you tell them what kinds of stuff Microsoft has/can put into a dossier. Nobody thinks: "Ah, Edge on a fresh Windows install, I'm glad Microsoft knows every site I visit, and can tell I'm a friend with someone because we use the same bluetooth speaker."
[0] It seems wrong to use the verb "leaks" when it's so obviously intentional.
If you use a different identifier to track changes in a GDID/machine-id/etc., that means you can continue tracking using mainly just the new machine-id, but you should always keep trying to correlate it with other things in case it changes.
But it would be very foolish for a blackhat to turn on browser sync...
Using the school news paper shared account, I copied the bullies coursework in to the public share, made changes to the coursework on my own user account and then copied it back to their work folder using the school news papers account.
MS Office keeps an "Last edited by" field so I got caught. If I had used the school newspaper account to make the edit I wouldn't of been caught. Rookie error. I too discovered a DCOM in the Windows 98 help file that revealed all the hidden shares on the network which scared the school. This was 2004 and I was 15.
Oh and the time, I bought a BB gun off someone at school. I was a librarian in the schools library and showed it off someone of the older grade year. The head librarian confiscated it and reported it to the headmaster. The guy then gave me a copy of Q3A and I stuck with violence the video-game way.
Not in the US, of course.
Many, many apps read /etc/machine-id if you do a quick github search.
Apps may have been silently correlating our activity for years without us knowing.
We know DHCP, EFI, GNOME, popularity-contest and many other apps already use it. There are countless ways it could be used already that are hard to detect.
It also helps that "Linux" isn't a monolith. One person's installation can be very different from another in terms of the software used. If one piece of software collects, that collection isn't as valuable as Microsoft's because the userbase is much smaller.
It's not perfect, but it's a hell of a lot better. A lot more resistant to abuse. Security in a lot of contexts tends to be relative like that. One's house isn't impenetrable; it's just better than another, in part because the neighborhood is better.
The number of new security bugs that e.g. LLM models have been finding lately, most of which are very old, are staggering.
Whether everyone does that correctly, that's a different question. On the other hand, while developers should be more careful about how they use identifiers like this, if an application on your machine wants to track you they don't need /etc/machine-id.
The issue with the microsoft account was not unique IDs - tons of unique things on a machine.... it was a unique thing tied to an account and shipped to remote places.
/var/lib/dbus/machine-id
Exists on Slackware64 15.0 but /etc/machine-id
Does not.To be clear, I think this whole story stinks but the issue here is not that OS installations have unique identifiers (machines have lots of unique identifiers -- hardware MAC addresses, peripheral composition and device IDs, screen resolution) -- it's that Windows has an auto-enabled spyware^Wtelemetry system that shares enough data for them to be able to provide that kind of data to law enforcement. This would be a privacy nightmare even in the absence of a unique OS installation identifier. Heck, they could even take regular screenshots of your screen but that would be too obvious[1].
However, by way of intentional code that goes against the interests of the user, which would be considered security issues with FOSS but not with proprietary codebases where the interests of the company come first, things should be better with FOSS.
The latter should also be more noticeable than the former. It's possible some of the former may not have run at all, neither accidentally nor intentionally. A bug can just stay dormant without ever surfacing in practice. The latter is there to be used.
The benefit of FOSS with respect to eyes is that those eyes are better aligned with the interests of users (since it's the users looking) than the eyes on closed-source code.
You can most likely get better security than mobile's, you just need to e.g. learn to write your own SELinux policies, etc. Facilities are there; they just have a learning curve.
Another option may be to set up a container or PID namespace and give your tool direct access to that /proc.
Regarding SELinux, looking at https://unix.stackexchange.com/questions/767564/selinux-deni...
> the entries under /proc/<pidnr>/ are running under the respective pid's domain
It also seems doable, since you can differentiate which PID directory belongs to what by the domain.
So I want to follow the principle of minimal privileges and only grant access to files needed for running a program. Sadly many programs cannot run without /proc. For example, poorly coded Apple's Grand Central Dispatch library crashes the application (for example, Telegram) if it cannot enumerate the information about threads or processes. It needs this information to calculate how many additional worker threads need to be created, and if it cannot calculate the number, it terminates the application for reasons I do not understand.
There is also a catch that the program can create a new unprivileged user namespace and mount /proc there thus bypassing my daemon completely. Anyone can create a user namespace nowadays.
> Another option may be to set up a container or PID namespace and give your tool direct access to that /proc.
/proc contains information not only about processes, but a lot of extra information.
If you want to completely remove the global information in /proc you can do so by mounting it with subset=pid (and extra points for hidepid=4). This was added to improve container runtime security but unfortunately a lot of programs still depend on reading global procfs files so we can't enable it by default.
As for why all this information exists in /proc, procfs has historically been a bit of a dumping ground. There is an argument for "everything is a file" and so on, but having had to deal with all sorts of insane bugs related to procfs I'm less sympathetic to that argument than I used to be.
> There is also a catch that the program can create a new unprivileged user namespace and mount /proc there thus bypassing my daemon completely. Anyone can create a user namespace nowadays.
There actually are some restrictions, if the process doesn't have access to an unmasked procfs mounting /proc will fail, though a recent patch[1] finally made it so that subset=pid works in this case. (And most sandboxes disallow unprivileged CLONE_NEWUSER with seccomp.)
[1]: https://lore.kernel.org/linux-fsdevel/cover.1626432185.git.l...
Also, I might want to provide some fake data, for example, if app relies on reading /proc/cmdline or /proc/cpuinfo. So the app should be able to read the files, but not the real data.
> And most sandboxes disallow unprivileged CLONE_NEWUSER with seccomp
I ma worried that some applications (like Chrome, Electron-based apps) might not work without user namespaces. GTK uses "glycin" library, it launches subprocesses for handling images in a bwrap-based sandbox, and it broke inside my sandbox, so I had to do quick fixes for it.
No, sandboxing in Linux is not easy. Just look at "man capabilities" or "man user_namespaces" and see how many rules and exceptions from rules are there. You need to understand it if you write a sandbox, but it is so complicated. And obviously AI cannot be trusted with such a responsible task.
[1] https://huggingface.co/blog/agent-intrusion-technical-timeli...
> How do you [...] allow reading only certain files in /proc, where process IDs are not known ahead?
Such things like /proc/cmdline should be even easier. You can just set regular DAC permissions and ownership on such files. You can just set ACLs on such files.
> There is also a catch that the program can create a new unprivileged user namespace and mount /proc there thus bypassing my daemon completely.
You don't get the ability to mount just because you created an unprivileged user namespace.
> For example, why do applications need to know the kernel command line? What for? Why do they need the list of major and minor device numbers? List of filesystems? Network configuration?
Because you might ask them to get that info. If you wouldn't have a use for that, you prevent that using one of a number of security mechanisms, including writing your own SELinux policy where you don't whitelist that access, setting DAC permissions and ownership as such, ACLs as such, etc.
Applications compose and are heavily configurable, and they include such things like language interpreters. You may want to write a script that uses any of that info for whatever. You may want e.g. your window manager to show that info on a statusbar periodically. That info can also be useful for conky or htop, etc.
Vim is a text editor, but you can e.g. have it make a call with ssh to get info from wherever and put that info into the buffer. You can say that a text editor doesn't have a need to read network configurations, but maybe by composing with ssh it does, and it does something useful with it.
I am not sure if I can set ACL or ownership on /proc files, it is not a disk-based filesystem.
> You don't get the ability to mount just because you created an unprivileged user namespace.
Please consult "man user_namespaces", section "Effect of capabilities within a user namespace". It clearly says that one can mount /proc inside a mount namespace governed by a user namespace without being root. The reason why you weren't aware is probably because the rules related to namespaces and capabilities are extremely complicated.
> Vim is a text editor, but you can e.g. have it make a call with ssh to get info from wherever and put that info into the buffer.
I definitely do not want proprietary applications running ssh on my system without me knowing.