Select * from Internet.blogposts(pfrazee.leaflet.pub) |
Select * from Internet.blogposts(pfrazee.leaflet.pub) |
There are multiple sites that support follower semantics over RSS. Feedland tracks subscriptions publicly, so you can see the blogs I read (https://feedland.com/?username=robalexdev), and who reads my blog (https://feedland.com/?feedurl=https%3A%2F%2Falexsci.com%2Fbl...).
I run another variant which collects OPML blogrolls via crawling, so you can find out who else likes your favorite blog and what else they recommend. Here's the page for Simon Willison's blog (https://blogroll-network.alexsci.com/discover/feed-a34ee2a88...). Thinking of RSS and blogrolls as a network feels much more resilient than blueskys Jetstream api endpoint.
Nothing to do with Bluesky services. There are many independent firehoses and relays. Here's a stream of standard.site blog posts coming in over a firehose hosted in Chennai. Every single one. No curator between me and the posts, and no work by me to crawl the whole network for them. https://pdsls.dev/jetstream?instance=wss%3A%2F%2Fchennai.fir...
Edit: ok there is an aggregator, the relay is scraping all of the PDSs out there to build the event stream. Notably this is not possible with RSS, where you need to build a large index with knowledge. PDSs request relays to crawl and that's that.
This stream lacks any blog posts that haven't been specifically published to atproto. Conversely, I see RSS feeds for the content in the stream, like https://bsky.app/profile/did:plc:byf7jvh3yvhffackiumpddtf/rs... (if RSS feeds exists for every profiles then I'd argue that atproto is a strict subset of RSS, but I'm not certain how that feature works). The stream may give you every blog post that was published to atproto, but that's already a small subset of long-form content on the web.
> Notably this is not possible with RSS,
The web supports blog discovery so well that people forget it exists. A simple Google search can find a very large amount of content, much available over RSS/Atom. Many web-based feed readers index the subscriptions across all their users to help with discovery. Here's the Feedland firehose, for example: https://feedland.com/?everything=true
I don't think its helpful to argue how complete the Google index is vs a bluesky firehose though. Most users are drowning in content and discovery is about search and filtering. There's lots of interest right now in vouching for the content of others: https://susam.net/wander/wander.js, https://codeberg.org/robida/human.json, https://www.manton.org/2024/03/11/recommendations-and-blogro..., etc.
One thing I'd love to see with rss / ompl sharing sites is more ease of exploration. It feels a bit clunky browsing many of these feed sharing sites because you need to evaluate each feed by manually clicking through each one etc or worse you can only import the whole opml at once flooding your feeds. For example I've seen some sites with feeds directly showing posts as they come in for each persons feed list(s) so you can see what that batch of rss feeds looks like in action without clicking around much. There are some sites doing great work getting people to share their lists but I think there is room for qol improvements.
One site I've been liking that has feeds for each list and is trying to reduce friction on sharing is blogflock.com. You can also follow other peoples lists directly on the site too so they show up in your main feed. I believe I've also seen some people self host their blogroll on their personal sites with similar feeds but I'm having trouble finding that software atm.
Absolutely. One good thing going for this approach is that anyone can grab the OPML export (the URL is stable) and build their own frontend, I'd love to see more.
Could be a fun weekend project for frontend folks.
You can add Reddit to that list, IPO is coming so better price out the third party apps we encouraged developers to build.
When Reddit raised its API prices in 2023 in order to make itself more attractive for an IPO, it effectively priced out third-party apps, such as the Apollo reader app. Discussion back then: https://www.reddit.com/r/apolloapp/comments/13ws4w3/had_a_ca...
- Semantic HTML
- Web indexing services
In this case, it's mostly indexing and a little of the other two if you want to add social features. I despite ATProto with all my soul for taking a problem with such a simple and standard solution and totally obscuring it behind hundreds of layers of JSON, faux federated services, and technical jargon, all in service of creating an inferior version of Twitter. I guess it wouldn't be as sexy to offer a web indexing service instead.
One can imagine making it a little more featureful, for instance indexing OPML blogrolls which would allow you to see whom a person follows and https://microformats.org/wiki/h-entry which would allow outgoing links to be categorised (like, reply, etc) so notifications filtered. With that, it would be possible to add a front-end site imitating one of the popular social media paradigms (Reddit-style or Twitter-style being the most obvious).
If there's one genuine design decision I would give for it beyond what is basically a cobbling together of existing interfaces, it would be to charge users per page uploaded. Likely very detrimental to growing the service, but I think one of the simplest ways to weed out spam and junk. A real problem with many web services is that the receiver of the message pays for it in terms of attention, where in other mediums the sender has to pay. Given that uploading crap is basically free, that's all you get. Increasing the cost of upload would weed out those endless AI summaries and lists of affiliate links.
The issue with ATProto is that it operates on too many layers. You can see lacking in what I have described here the concept of durable authorship, but this is a property of content not how the content is distributed. Perhaps someone will invent a standard way to sign HTML documents, in which case you could base user accounts on that instead of DNS. AT enforces this centrally but it does not need to. It is walling itself off from the common and decentralised software ecosystem of the web for no good reason.
It'd be great if schools taught foundations like how the internet works (not in-depth, but basics like what a server is) and how to create your own server. People would get gradually less afraid of technology and that'd enable systems like this.
As it stands, people are just more comfortable using walled gardens because their data is managed for them and they don't bear any responsibility for it.
If decentralization is ever going to win, it needs to be turnkey, explained without showing a network diagram, be basically free to run on your laptop, and actually have the content people want to see and not just be a bunch of people who speak lojban. [1]
While I'm not sure how much of a legal leg X has to stand on, I understand why they'd rather Nitter not exist. Twitter tried (valiantly, imo) to stay open. What ultimately began the API lockdown was the need to stop bleeding financially. Ads were inevitable, and there being no reliable way to do that via API access.
Twitter solved those three problems, and no one spends nearly as much time talking or caring about the technology used to do it than people who talk about decentralization. Myself included in my younger days. Now they're just defending their moat.
You can choose where your PDS runs and the ecosystem still works. If one relay shuts down (Google Reader) or takes their API private (X, Reddit) it shouldn't matter, the PDS are separate and another relay can take over.
Worst case you can go directly to the PDS, but that wouldn't work at scale
Oh! So you don't know very much! Well, happy to help give you some links so you can learn some of the basics!
@pfrazee gave a talk recently on how atproto was built from day 1 for scale. It's not a secret, it's in the name, Authenticate Transfer Protocol! Your key that you control signs all of your content! Such that your can be forwarded by anyone, that it continues to bear your authentication, wherever it goes, whatever replicas it's on!
https://youtu.be/BoJnj2yPf14?is=YcnoJZSqQw1YBhl9
It's so easy and small that one can run a whole network firebose off a small VPS! There's a list of them. https://atproto.at/relays
Discoverability is a function of this! There's all sorts of tools to filter for different kinds of data. There's indexes, so you can find things like back links, https://www.microcosm.blue/. And there's specific schemas (lexicons) for helping find blog posts, https://standard.site.
I'm a fan of the API that Solid exposes, I'm a fan of the json-ld. But for discoverability and scalability, in terms of being a networked protocol, ATproto is great.
Fortunately this isn't true. Social media consists of a few big walled gardens, but the web itself is still open. Go forth and create a website!
So the question is, are you okay with giving random people a mirror of your public posts? After all, they're public. It's like putting them in a repo on Github for anyone to clone.
Google's Go Module Mirror is a similar but more specialized service that has a copy of all the Go modules that anyone has published on the Internet.
That makes no sense - and yes, I'm aware that people think like this, even though it's entirely contradictory.
If you intend to make information public, you relinquish control over access to it. Trying to walk that back, or even to retain capability of walking it back, is perverting the system and turning the earlier claim of information being public into a lie.
i think the atproto has mechanisms for this?
I agree it would be neat.
But, when did we conclude this "should" be possible and thus warrant such a whiny post?
We made it a pretty big goal from the start to clearly communicate to users that the network is extremely public, and that we believe open access to public content is an important part of preventing another round of walled gardens. This is generally understood and appreciated, but the community has a wide range of opinions about it. Some people see public as public and that they’re there to have their voice heard. Some people see it as a pretty uneasy arrangement at best, and would rather it wasn’t that way.
The atproto community has been spending a large portion of the year developing “atproto spaces,” which are essentially a way to cheaply mint mini-atprotos with access control. This is the answer to non public data. It will be added to the Bluesky app and the ecosystem, and it will remain accessible to apps based on the grants of the users, but the spaces will not broadcast their content on the firehose. I believe, based on the reception we’ve received from that work, that people appreciate that it’s being done, and that people eager for less exposure will pick it up.
This should, hopefully, resolve any tensions at play. We will continue to advocate for public speech and I think many people will continue to participate in it, but I am quite curious to see if the non public spaces gain more adoption overall. My guess, based on people’s general use of the internet, is that they naturally will. The public arenas serve a particular purpose and not everyone is always trying to be a part of them.
EDIT: did I really just say "let me give an honest accounting of that." I have been spending way too much time with llms.
I don't think that's true. For example, nowhere on the blocking interface does it tell users that their blocks are public.
My guess, based on what happened in the past decade, is that it will definitely happen, a lot of people will be unhappy about it, but by then it'll be too late to stop it.
Private spaces have, by definition, a natural asymmetry to them: it's easy to find a plausible reason to turn a community private, an argument that resonates deeply with people, even if it's not applicable or completely bogus. It's hard to convince people to appreciate the fact that conversations in the open have value reaching beyond the participant. Public spaces are a gift to the world.
The way I see it: imagine WhatsApp, et al. came a decade earlier, and the Internet went dark in the early 2000s, suppressing the "blogging revolution". The idea of Internet as source of knowledge, a place where you can learn anything you're curious about, would've never happened earlier, as all knowledge would've gotten locked in private groups, out of reach of search engines. That's what happening right now - even in OSS, if not due to private groups then due to adoption of Discord and Slack as primary communication platforms. The public Internet is increasingly just commercial slop. But it never would've been anything else than slop if private groups came before blogs and social media.
Decentralized identity. You don’t have to choose a specific instance and be locked into your decision.
ATProto is attempting the "if we make it P2P it'll be popular!" strategy except not with P2P.
There's a number of different tools to find accounts out there. The main way apps do account type-ahead is a free 3rd party service, one you could run yourself. No one does non-centralized because the network is easy to index & there's no use case for trying to ask everyone each time to do the work, that's a bad strategy, why do it?
Kinda feel like you are just blowing chaff, moving goalposts, not interested in finding out or learning. So, yeah, seems like we're done.
The underlying tension is the "dark forest". Opening yourself up gets a bigger audience. That feels like a good thing, until your reach expands to those hostile to you. Every community has always had to deal with abuse, but now there's much more coordinated, intentional attacks on communities, orchestrated (ironically) over social media.
A) since everything is public it's much easier to see what someone is about, what they are doing. Including the people coming to disrupt and inflame.
B) Bluesky's blocking actually gives folks the ability to deal with their haters, to block reposts for example. So folks can manage engagement they don't want.
It feels like there's almost no cases where communities are given the powers and capabilities to manage the problems. The attackers have been given all the benefits of dark forest advantages, with none of the downside. It feels like there is finally a bit more of a symmetry, and some systems for defense, which until now, no social network has allowed its users.
It's far from over, who knows how things go. But so far the antagonist world burners have not had nearly as good a chance setting fire to the scene as they have had everywhere else. I remain cautiously tentatively hopeful.
A way around that is to use a pseudonym, but it requires care to not post anything revealing.
I think it is generally bad to implement new features like this on the part of the indexer. It should just keep track of an existing web of documents rather than creating it's own format and walled garden.
Ah right, of course. Fair enough. I think Dave Winer is trying some similar ideas, though I haven't looked at it beyond knowing he's trying to extend RSS to support these capabilities.
Why, instead of creating an indexer that works with existing formats, did they choose to create one intentionally incompatible with 99.999% of web content? Because in practice it's not a web protocol, it's not a serious thing. It's just a whole bunch of bullshit for nerds to make blog posts about that serves as the back end of a Twitter clone. It is standards proliferation. It adds nothing and fragments existing efforts. Instead of contributing to the web ecosystem, it subtracts from it. It breaks the important rules of software development, to do one thing and do it well, to use existing solutions. Everything it does is duplicated, there's a new way to sign documents, a new way to host content, a new data format, a new exchange protocol, a new way to serve documents, a new way to send notifications, etc. It's not a tenable way of doing things. If every person looking to make a web indexer created an entire separate web ecosystem, that would not be sustainable. They're clearly not intending to replace the Web, so what are they doing? It's just a complicated way to store tweets.
And the silliest part is that, if I recall correctly, they don't even have decentralised network yet since only they can issue PLC DIDs.
For what it's worth, the W3C recommends ActivityPub (fediverse base protocol). Do you like that one or do you only like the HTML stack?