Yes, yes, we should fight deceptive software that doesn't disclose it has a proxy function. Fine. But there's no way to square 1) fighting all residential proxies, including ones run with informed consent and 2) preserving the public's access to general-purpose compute at home.
I'd rather live with the proxies than for someone to tell me that if I run squid, I'm damaging national security.
When it comes to involunary proxies, those are really just one of many possible symptoms stemming from a real problem: Shitty security. (Edit: And shitty contract/privacy laws.)
Shitty security is tolerated by our markets, is is protected from fixes due to copyright law, and it is even encouraged by parts of our government that want to exploit the flaws. We will gain far more from fundamental quality-improvements than we will from whack-a-mole-ing on this symptom.
> We will gain far more from fundamental quality-improvements than we will from whack-a-mole-ing on this symptom.
100%, but I feel like that's both true and rarely executed upon in most governments for most topics.
They are basically 'won't someone please think of the children?', but with higher stakes.
The abuse of the term "National Security Threat' doesn't mean legitimate threats don't exist. "There are no NSTs" as a left-wing dog whistle would be equally damaging by ignoring legitimate concerns - both approaches lead to the same outcomes, similar to the vast majority of highly popularized left-right wrestling shows.
Hidden, unknown to the public residential proxies are a national security threat, and there's no need to ban them in order to protect from that - just make them public. There would be no problem if I knew the IPs who could act as proxies and block them, that would be a fair game.
It strongly depends on how scraped data is being used.
If your 'cat' hasn't been able to catch their 'mouse' then the cat needs to get smarter, or look for alternative sources of mice, or the cat should be considered 'unviable'.
Have you approached them to get access to their data? Have you explained to them how your service can benefit their business?
(I generally come from a position of suspicion as to why someone wants to scrape data that the owner goes to certain lengths to protect, but then I'm also an 'information wants to be free' kinda person, but the Internet is increasingly an untrustworthy place, so security is overruling narrative).
Let's try this for example. Suppose you want to create a price comparison site.
A lot of the major retailers don't want these, because they want the customer going to their site when they want to buy something, not to the price comparison site that tells the customer which retailer has the best price and then might not always be them.
If the comparison site doesn't have the prices from retailers who sometimes have the best price then it can't serve its function -- customers still have to check sites manually. And comparison sites are pro-consumer whereas blocking them is anti-competitive, so the ones doing the scraping are the good guys.
Specific Domains:
proxyjs.brdtnet.com
proxyjs.luminatinet.com
proxyjs.bright-sdk.com
clientsdk.bright-sdk.com
clientsdk.brdtnet.com
Wildcard domains: *.brdtnet.com
*.luminatinet.com
*.luminati.io
Source:
https://blog.includesecurity.com/2026/06/the-smart-tv-in-you...The overall residential proxy market is too large and often undetected by intelligence tools. BrightData is actually much more compliant and malicious actors wouldn't be allowed access. They have an extensive KYC and use-case vetting process.
And as already pointed out... many of the more popular networks actually do have KYC processes in place, in addition to blocking specific targets from being scraped through their network(s).
So a pretty good first step in blocking then.
And also assuming that the only reason for higher latency is physical distance rather than crappy WiFi or corporate nanny filters or the client device swapping because the user can't afford more RAM.
> Actual data and identity loss of American citizens.
> Malware that can record video and audio from infected devices.
> Infrastructure for foreign covert influence campaigns.
> Botnets used in hacking and DoS attacks
These things have existed for 20+ years. Bad but not exactly a national security threat.
People were dismissive of security far too long. I think it's actually good that AI instills some fear into people causing it to take a lot more seriously than they have previously done so.
Sometimes their system malfunctions (cough cough) and you get charged for those clicks but you fail to report them as fraudulent as you've no proof as the IPs and useragent all appear normal to ad agencies and advertisers.
Proxying is least of the worries!
When manufacturers bring out the smart features that require constant data collection, monitoring (obviously in the name of service quality, troubleshooting, firmware / app updates and subscription of services (!!)), it is oblivious to the fact that security and privacy at a personal / state / national level is never considered (for various reasons such as the design, profit interests and overheads / compliance perspectives and what not!) to be part of the ecosystem. Bad actors with sufficient knowledge and tools could exploit and profit from it.
It has become more of a planned obsolescence / profit / greed based industrial / corporate culture in rolling out unwanted stuff / monitoring and controlling as a feature that gets exploited to the core.
Honesntly, I don't have an answer for that, but have some thoughts that I wanted to share.
I do not want a cleaning robot or a water purifier or a printer or a TV or a refrigerator or an oven or a smart light bulb or a smart lock or even my car to have undisclosed, unwarranted connectivity to internet for whatever reasons.
I do not care if it is the manufacture or the service provider or whomsoever it may be!
I want all of the network connectivity options to be explained with all the controlling options and boundaries and data collection that happens on any of the connected medium to be as much transparent with the option of preventing or restricting what one does not want to go out of the device.
Will anyone do that?
It ill never happen! Sadly!
This problem will balloon further without adequate controls and killing the networking at a ground level would only be the only solution!
But, in the case of a connected device having an inbuilt connectivity option (like the m2m based options) in a smart vehicle, even that is not possible!
Continuing to allow CGNAT is what allows the residential proxies to hide because it tumbles the IP identity of bad actors with normal people.
The nation's security infra is a reflection of the nation's security legislation and regulation.
Aside: TheSInIoT is my DMZ's wifi password.
More importantly, a residential proxy can be mutual. If a group of people agree to forward traffic to each others’ IP, how is it not their business? It obfuscates their identities, but this is the Internet, not e.g. a government office or test center where you obviously can’t walk in with someone else’s ID.
Banning non-consensual proxies makes sense since those are effectively malware, even though they make anonymity harder, so does banning theft and here you’re stealing someone’s internet. I’m sure we can convince enough laymen to knowingly install proxies by paying them.
Myself being "not a layman" would definitely not allow anonymous usage of Internet that's tied to my home/residence/identity to be used by some internet rando just because they pay me money.
I would support, however, small family/friends groups who know and trust each other well enough to not do anything that'll cause a police raid on each others homes (having endured a police raid and the ensuing 8 months of not being told anything about the progress of whatever they're doing, it's not something I'd wish on a stranger, never mind a friend/family member).
Sure, the service should be upfront, although most laymen probably won't read. But if people had to really understand consequences before they accepted anything, most would miss out on most overall mutual offers. Importantly, the expected consequences are low (as long as the service isn't greedy), and the ratio of tech people agreeing would suggest that if most laymen did understand they would still.
The interviewee seems pretty inept yet still managed to uncover something really interesting and nefarious.
Can’t wait to have our very own Roskomnadzor MITMing everything and national security laws mandating Kazakhstan style TLS interception on every single device.
Just a small over the air update via your friendly neighbourhood corporations Microsoft, Apple and Google.
You know. To keep you safe.
> At the very least, major American ISPs (Comcast, AT&T) should detect clearly suspicious activity coming from customer IPs and warn them to scan their computers, check their TV apps, and find whatever is turning their internet into a proxy.
What are we saying? Really, what are we saying? We should turn domestic ISPs into domestic surveillance apparatuses to "detect clearly suspicious activity"? What is "clearly suspicious activity"? Section 230 is still law.
Governments could even purchase residential proxies themselves and check whether they're being used for malicious activity.
It isn't that difficult, but I guess governments just aren't that interested.
> It isn't that difficult, but I guess governments just aren't that interested.
I recently switched from my previous ISP because they randomly broke my ability to use my own router, and during the time when I had to default to using their own modem/router hardware before I could get the new ISP to come and set things up, I could barely even get a signal in my office upstairs (which had previously been connected via a mesh endpoint) because their router didn't expose any way for me to split 5 GHz and 2.4 GHz, and the router absolutely refused to let my devices connect via 2.4 GHz despite them having more than 90% packet loss due to the weak 5 GHz signal.
Regardless of how "easy" it is, I don't trust ISPs not to screw it up somehow and probably cause a lot more concrete damage (even if the individual issues they cause are smaller in magnitude) than the theoretical concerns of "national security" that, as far as I can tell from reading this thread, have caused a total of like a few hours of downtime one day in a couple decades.
Besides, mobile proxies exist which work differently. All you need is a mobile data plan that you run a proxy server on. It allows for easy IP rotation and since it’s a pool shared with other customers you cannot easily block it. This is because of things like CGNAT.
IP rotation is easy because reconnecting to the network gives you a new IP address.
Static residential proxies also are a thing, even if they are less effective sometimes.
Separately, would you let your home/personal connection be used as a residential proxy?
Personally, I would not, except if it was for a friend / family member, but I would still ask them, pointedly, why my internet connection is required rather than their own.
No one is being hurt by someone sharing their internet with me a few times a week.
Just because the idea of net neutrality exists doesn't mean it's true. I'm sure there's a specific context for it, and 'security' is not that context. It was about data/packet prioritisation wasn't it? Unrelated to security.
lets just hope some day there will be enough threats for nations to crumble
Do you allow the data you've scraped to be scraped? Do you share it as freely as you desire the 'companies that want to hide it' would? Or do you consider the scraped data is 'hard earned reward for effort' and therefore has value that others should subscribe to your service for?
> you need a video call and KYC to get access to a high quality IP pool
They need to know who to blame (cough come after cough) if their top-shelf stock gets tainted by some script kiddie.
The I guess the site is either useless, or can be used in combination with customer's own research that includes sites not included in the comparison. The site is still useful, it's just not exhaustive. There's nothing much in the world that's exhaustive, it's all a range of percentages.
Also, maybe the source that isn't scrape-able becomes less popular as a result of not being included in the price comparison list. If they're always more expensive, then neither you nor they have any advantage in listing on your site. If they're always cheaper, then your site may have no purpose to serve.
Everyone seems to want _everything_ despite _everything_ not being necessary to provide a service. The lack of a certain set of data may itself be a data point, and should be used as marketing material for or against the company resisting the scraping.
Just because residential proxies are somewhat of an 'easy answer' doesn't mean they're right. Use more creativity and imagination! Or follow a new idea. It's not like price comparison sites are a revolutionary idea.
(I'm a consumer who happily does significant research before a buying decision, going to various sites and checking prices, warranties, model numbers, reviews, availability, delivery times, and all the shit. And prefer doing the research myself than trusting a price comparison site, so I'm not the target market, thus my bias in my commentary and opinion on that topic).
> so the ones doing the scraping are the good guys.
Everyone thinks they're the good guys. If there's profit involved then there's always an element of delusion.
There is a significant difference between having the prices for 85% of retailers because you haven't figured out how to get pricing from 15% of them and having the prices for 5% of retailers because the biggest ones block you from getting their prices. The amount of work you can save the consumer, and thereby the usefulness of the service to them, changes by a dramatic amount.
> Also, maybe the source that isn't scrape-able becomes less popular as a result of not being included in the price comparison list. If they're always more expensive, then neither you nor they have any advantage in listing on your site. If they're always cheaper, then your site may have no purpose to serve.
Now consider the possibility that they sometimes have the best price and sometimes don't.
> The lack of a certain set of data may itself be a data point, and should be used as marketing material for or against the company resisting the scraping.
Suppose you go to the price comparison site to look for the best price, the best price in its data set is $22, then you go to Amazon and they have it for $19 because it doesn't have their prices. The customer then starts checking Amazon in addition to the price comparison site. Their price is only actually lower 15% of the time, but another 75% of the time it's exactly the same, so the customers who don't want to keep checking two sites start checking only Amazon instead of only the price comparison site, at which point Amazon gets to charge them a higher price on 10% of stuff, or a higher price on more than 10% once people have been trained not to check. And prevent them from patronizing a random competitor when their price is exactly the same
Causing that to happen is the reason they don't want their prices in the comparison site. If not being listed there hurt them then they wouldn't be trying to prevent scraping.
> I'm a consumer who happily does significant research before a buying decision, going to various sites and checking prices, warranties, model numbers, reviews, availability, delivery times, and all the shit. And prefer doing the research myself than trusting a price comparison site, so I'm not the target market, thus my bias in my commentary and opinion on that topic
Consider also that you may be able to justify doing this when buying electronics, or choosing which brand of something to buy on a recurring basis, but if you just want the best price on the product you already know you want, it's crazy to spend hours to save cents. But completely sensible to switch to a search box that would save you 10 cents on every $2 purchase by giving you price-sorted results from multiple retailers who all sell the same products, if it's allowed to exist.
Do we have solid numbers for that?
Trying to restrict things to end-users is the total opposite of that. The common addresses everybody gets automatically are the thing you're trying to allow. Address reputation is pointless because residential customers get dynamic IPs and the reputation you're trying to record for some IP address can get swapped with a different customer at any time.
It also feels like a description of the perfect camouflage to facilitate doing bad things: "don't block them because you might block an innocent bystander". Putting innocent bystanders in harms way sounds like someone else is the bad guy, not the person doing the blocking.
The problem being that attack traffic is disproportionately coming from devices that are compromised, which is already illegal, and you can't fix that by making it harder to use residential proxies for things that are legitimate, like sharing IP addresses between real users so they can't be used as a personal tracking ID.
> It also feels like a description of the perfect camouflage to facilitate doing bad things: "don't block them because you might block an innocent bystander".
Cloudflare promotes putting your site behind Cloudflare to inhibit censorship, because then the censors have to block their entire service (which is half the internet) to block anything. It's not always a bad thing.
> Putting innocent bystanders in harms way sounds like someone else is the bad guy, not the person doing the blocking.
If there is an alleged thief on the subway and you respond by lobbing a grenade into the subway, there is more than one bad guy.
If it has to be scraped then there may be other problems (which include lack of resources to make the data API accessible).
That's cutting the tether from a _lot_ (there's no way to overstate this) of useful information, but that's potentially one of the great things about the open AI/LLM models, is that all that info is baked in there.
Disclaimer: As far as I understand it. Please educate me if I'm way off the mark.
Also, how much of the content of reddit is in Common Crawl? (same disclaimer applies to this comment)
I also understand 'the archival mindset', I hoard a bunch of data. But I also understand the logarithmic graph of futility.