Oxide Computer: Docs(docs.oxide.computer) |
Oxide Computer: Docs(docs.oxide.computer) |
https://oxide.computer/podcasts/on-the-metal
I just noticed they have a second podcast. I'll assume it's just as good.
[0] https://oxide-and-friends.transistor.fm/episodes/the-fronten...
This is not a dig on the program at all; I'm glad they are making the time to produce it, and I'd rather they spend their effort getting racks out the door instead of generating marketing hype.
I did a little work in the late 90s with Alpha-based machines. I was impressed at those machines didn't seem like the hack-job crap that PC-based stuff was (with simulated chips from the early 1980's hiding out in dark corners because "compatibility") and still is today. I'm betting working with Sun gear felt similar, though I never got to work with it. Just having an honest-to-God serial console, as opposed to crappy bag-on-the-side things that scrape video memory and pretend to be "legacy" PC input devices, would be an amazing thing.
I'll never be able to work with their stuff because I don't work with Customers at that scale. I'm also vastly unqualified to work for them at their current stage. I suppose maybe someday they'll need field service technicians... I can hope, I guess.
(Sorry to be that guy, but just a friendly suggestion: high contrast dark themes are difficult to read for people with astigmatism. Especially since this is technical documentation, intended to be thoroughly read, you might want to consider a light theme toggle.)
Second to that I just want to say the presentation of these docs is top notch. (I so desperately wish I was the target customer for these systems; reading these docs makes me want to do terrible things to my electrical service and play with one of these racks.)
On-prem servers aren't a new invention. The market seems pretty saturated (and shrinking). Virtualization isn't a new invention either. The market seems pretty well served, at least commercially. Can the integration of both be a convincing enough advantage?
The management UI certainly looks nice; it's something I'd like to have on my KVM box at home (any good Proxmox alternatives?). I don't see why it'd have to be bound to an enormous server.
I would put Wireguard/Tailscale in the same category as well.
Is this just more of the same, or is there some innovation there? It is interesting that computing trends go in centralise/decentralise cycles. We can observe this as far back as 1980s. I can't wait for the next "decentralise" cycle as I'm under an impression a lot more innovation happens during that phase.
The innovation is mostly in the integration between the hardware and software of both the server and the switch.
This podcast goes into what that enables: https://www.youtube.com/watch?v=AkWh2Sms3aw
> I can't wait for the next "decentralise" cycle as I'm under an impression a lot more innovation happens during that phase.
I think we are beyond that cycle. Centralization and decentralization happen at the same time. We have much more much cheaper chips now everywhere that do a lot more, and then we have local and distributed compute deployed everywhere and we also have large datacenters at the center of it.
Neiter datacenters nor distributed compute is gone go away anytime soon.
Can you elaborate on the Sun comparison? I am a huge fan of Sun and what they did for computing at large - designing hardware, creating specs, their contributions to evolving unix and so on. I'm not sure how Oxide compares. Unless you're talking about "in the spirit of Sun".
Bryan worked at Sun where he helped create DTrace.
After the Oracle acquisition, he left to join Joyent as VP of Engineering and then CTO. Steve was COO of Joyent. And I have heard similar comparisons between Joyent and Sun.
So is it 1 or is it 0..
https://www.siliconmechanics.com/system/rackform-a235.v8.1/2...
I think I will go back to my little collection of RPi Compute Module machines. And maybe some day in the very distant future I can buy big boy servers lol
[1]: https://docs.oxide.computer/guides/creating-and-sharing-imag...
Are there intentions to go smaller than this?
(my personal perspective as an Oxide employee)
Not in the near future, when a lot of the design relies on the benefits of scale, scaling down kind of removes some of the benefits. But also just like, things are still early, there's not a lot of choice period.
That said, we do have a lot of fans who want to buy something, and it would be cool to figure out how to do that. But also, we gotta like, make and ship the primary product. So we'll see.
How would an Oxide computer systems look like for for an "OpenAI super computer" compared to Azure?
https://docs.oxide.computer/guides/system/rack-installation-...
The rack also comes with built-in security features to ensure all hardware and software are genuine Oxide products:
purpose-built hardware root of trust (RoT) – present on every Oxide server and switch – cryptographically validates that its own firmware is genuine and unmodified
encryption of data at rest via internal key management system built on the RoT
trust quorum establishment at boot time to ensure the cryptographically-derived rack secret is verified before unlocking storage
https://docs.oxide.computer/release-notes/system/1-0-0I think the sleds blind mate to a 55V DC common bus.
You can certainly install Ubuntu on a very powerful machine with a WAN interface (e.g., a NIC connected to a residential cable internet connection). Then, use something like k0s to provision other bare metal workers. Those three steps, and you've got a "bare metal container platform." You don't need a "LoadBalancer", you can specify that nginx-controller runs on the host network of specifically the machine with the WAN interface and configure its service's external IP to the WAN IP, and now you support Ingress.
But how do you imagine having multiple LoadBalancer resources without multiple IPs? And how do you imagine having multiple IPs without ARIN? The turnkey challenging part is the public IPv4 addresses, not the platform.
And if you need to service that machine or it goes down?
> But how do you imagine having multiple LoadBalancer resources without multiple IPs? And how do you imagine having multiple IPs without ARIN?
Most colos/transit providers will happily lend you their IPs
See the excellent "Oxide at Home" blog post [1], and HN discussion [2]
[1] https://artemis.sh/2022/03/14/propolis-oxide-at-home-pt1.htm... [2] https://news.ycombinator.com/item?id=30671447
I like the idea, though it may be dumb, if you could fill it with different types of processors for what you need.
Sort of like chiplets but at a different scale
So "they probably really really -want- to show it off, but are showing remarkable restraint for good reason" seems at the very least plausible to me.
I had a chance to talk to someone who worked at Oxide once. I got excited because my experience matches up nicely with their products on several levels.
The person got excited as well and started talking about how to interview there, but the conversation got kind of weird. They kept emphasizing how important it was to not talk about compensation in the interview. Apparently they pay engineers all the same comp (not bad, but would have been a significant step down from every offer I received during that time of my life) and they select for people who aren't interested in getting paid a lot for their unique skills.
Probably not a bad deal for people who like working in that domain with like minded people. At the time I got some very uneasy feelings from being aggressively coached to not bring up compensation or ask any questions about equity during the interview like it was some unspoken rule that would get me disqualified. Maybe the person was exaggerating, but I found at least one other person with a similar story.
Honestly, their comp would have been awesome if I was a single guy living in a low cost of living location and working remote, but at the time it would have meant giving up quite a bit to work for an early stage startup with high expectations and an unspoken rule that I should never ask about compensation.
[0] https://docs.google.com/document/d/1Xtofg-fMQfZoq8Y3oSAKjEgD...
[1] https://oxide.computer/blog/compensation-as-a-reflection-of-...
Yes, that is a sad point. I love Oxide and everything about them. It is the only company in Silicon Valley that I'm honestly kid-like excited about (I fake the "passion" for others but don't feel it). My partner is probably tired of hearing me drool about this one company I wish I could work for...
But with a family to support, it's never going to happen. The pay cut would be brutal, so I never apply. If I ever become an independently wealthy multimillionaire, the first thing I'll do it apply to Oxide. As long as I need a paycheck, it's impossible.
Commodity servers have the same specs.
Here’s a reseller of SuperMicro servers, where you can buy similar compute on the cheap.
The thick layer of hardware contrivances necessary to maintain IBM PC compatibility is unnecessary for the task of bulk hosting of x86/x64 VMs. There's a lot of hardware and software that just doesn't need to be there.
Bare metal out-of-band management ends up being bolted-on to these "legacy" contrivances (scraping video memory for remote consoles, faking being USB peripherals). A serial console or SSH connection to a service processor would be vastly superior. I can't begin to count how many times an iDRAC "lied" to me about issues with a machine, or how many times the solution was "upgrade the iDRAC firmware and reboot it".
I have been mostly unimpressed with the quality of firmware for motherboards, baseboard management controllers, RAID adapters, NICs, HBAs, power supplies, backplanes, front panels, etc. Every new model of system or component ends up being an exercise in fear / anticipation of problems. The integrator has very little power over the firmware quality and I can be assured that if I do have a firmware-induced issue I'm many, many steps away from actually communicating with somebody who can help.
Granted, maybe if I was buying at the scale of Oxide's prospective Customers I'd have some pull with the integrators, but I'm skeptical of that, even.
Oxide is actually building computers. Putting commodity motherboards into boxes with other commodity components won't ever have the level of attention to detail and integration that Oxide can provide.
Probably not a coincidence. It would be interesting to know which ODM they partnered with for the hardware.
I've done some work with SuperMicro in the past. Some of their boards come with extensive headers and customization options right out of the box. They're also happy to work on board level customizations with the right contracts in place.
(Or maybe I speak for myself only...)
Thanks!
Darkreader is a lovely plugin that might make your life much easier: https://addons.mozilla.org/en-US/firefox/addon/darkreader/
I believe there is a Chrome extension, too.
It lets you set BG, FG, sepia, total contrast, etc. It is quite a neat piece of kit.
Are you yourself affected by this? I have astigmatism and I keep hearing this, but I've never experienced it. If you are affected, do you keep your screen at high brightness? I'm wondering if it doesn't happen to me only because my astigmatism is mild, or if the fact that I tend/have to keep my screens at relatively low brightness plays a role.
Incidentally, for different reasons, high contrast dark themes can be problematic for me as well (especially at high brightness). Dark Reader and Midnight Lizard are essential for me in keeping contrast in a comfortable range.
They go into more details on their podcast, and this section in particular covers the bootstrap time: https://youtu.be/5P5Mk_IggE0?t=3381 Pretty fascinating stuff.
I can't imagine that large cloud / "web scale" companies would want that. Most want a fair bit of control of their own hypervisor and management stacks based around KVM. And "enterprise" type companies are going to have issues with certification I would have thought -- will RedHat, Microsoft, SAP, Oracle, etc certify their supported products on top of this hypervisor? Seems like a difficult and expensive process.
So what's left? Companies that support their own virtual machine software but don't support their own hypervisor and don't like what's available from vmware or Microsoft or RedHat. A small niche. Or are my assumptions wrong?
Basically if you are on-prem, and you are dissatisfied with what you are getting out of today's onprem sellers. Things like bad firmware with slow update cycles, issues with rack/power supplies/cabling/interconnecting systems. Closed down systems that don't allow much customization, etc. They are open sourcing a lot of their work along the way
Again, I'm not affiliated with them, and my info may be outdated so take it with a grain of salt. But that's how I've seen them for some time.
I'd recommend the O+F ep posted in sibling, but I think here the pitch is "well you need this hardware anyways right? How about buying one that's easy to use and doesn't take a month to get working?" All built by people who are so obsessed with root cause analysis that they've ended up writing their own firmware, running on an OS where these people are common contributors.
There is an existing market, it is large. Making a product for that market that is just simply 'really good' has potential to make money.
Almost all great computer companies started into markets where things already existed that were comparable.
1. It's not clear the market case closes, because while there are customers who would greatly benefit from their systems, those same customers are also very averse to change, meaning that most of them will probably sit out for the first few generations to see if they have staying power. If everyone does that, it becomes a self-fulfilling prophecy.
2. ... but this would have been fine, a few years ago. It's not fine in the current market situation, where there is much less easy money looking for a place to go. I sure hope they are sitting on a long runway.
3. Speaking of, they are very late in their execution. A SM4 platform that only shipped in 2023 is not great. I really hope they are far along on their SP5 development. (... But this also dampens current sales. I bet a lot of potential customers are thinking that a SP5 Oxide Rack seems a much more appealing than a SP4 one, so why not wait?)
But the reason they are not discussed that much is that people who understand the market are all really, really hoping that they pull it off. Because the current situation in server hardware is dire. In every thread people who don't know much about servers ask why not get a similar system from Dell or Supermicro or whoever. The answer is that the commodity servers are pieces of shit that requires significant local engineering resources to manage. Software quality of firmware is generally horrendous, and when this is pointed out to the vendors, their answer is that they know, but everyone else is just as bad, wontfix. A significant draw of the cloud is that all that pain goes away, because the hyperscalers realized how shit everything was and fixed it in their systems. It just never trickled down to the market below them.
Sometimes the hype exists for a good reason - as a paying enterprise Tailscale customer, I’ll fight you if you ever suggest I have to staff people to admin an IPsec or SSL-based VPN again.
Other then that, the code is open source. They also help sponsor both Rust and OpenFireware conferences.
This is pretty darn close to the "Gimlet" Compute Sleds
One 225W TDP 64-core AMD Milan CPU 1 TiB of DRAM across 16 DIMMs 12 front-facing hot-swappable PCIe Gen 4 U.2 storage devices 2 internal M.2 devices 2 ports of 100 GbE networking
Point for point, the same hardware as specified.
If you have a team that is struggling building an internal cloud with all this control (and problems) and all this commodity hardware (and its problems) then maybe they would be happy to switch to something that just works.
> And "enterprise" type companies are going to have issues with certification I would have thought -- will RedHat, Microsoft, SAP, Oracle, etc certify their supported products on top of this hypervisor? Seems like a difficult and expensive process.
If that was the case and nobody running any of these would run their rack, then I wouldn't think they would not have received any funding. But I don't know enough about these certification process to really comment.
> Companies that support their own virtual machine software but don't support their own hypervisor and don't like what's available from vmware or Microsoft or RedHat
Non of these come with a fully integrated rack.
The competition would be somebody willing to buy a rack of Dell servers with VMWare software. Or somebody willing to buy a rack of Dell server and then use RedHat and set up all their own cloud style infrastructure.
That's not what I'm assuming here. Read carefully, I divide the market into 3 categories. Those who support their own VM image software and hypervisors, those who support neither, and those who support VM image but not hypervisor.
First is Amazon, Google, Facebook and the like (and it's not an assumption we can see their public contributions to KVM, QEMU, etc., and hear their talks about some of what they use internally). Second is "enterprise" who wants something that just works. Third is ? and would they want to support their software on a niche hypervisor?
> If that was the case and nobody running any of these would run their rack, then I wouldn't think they would not have received any funding. But I don't know enough about these certification process to really comment.
Well it is the case that enterprise (supported) software is not just supported on any hypervisor. https://access.redhat.com/articles/973163 RHEL runs on their own KVM as well as MS, VMware, some cloud vendors. Some application software also gets certified to hardware and hypervisors, not just operating system (e.g., SAP does this).
> Non of these come with a fully integrated rack.
It's not fully integrated if it doesn't come with the guest software though, is it?
> The competition would be somebody willing to buy a rack of Dell servers with VMWare software. Or somebody willing to buy a rack of Dell server and then use RedHat and set up all their own cloud style infrastructure.
Right. And the problem for Oxide is that the competition will have fully certified and supported operating system and application software for their virtual machines.
[0] https://oxide-and-friends.transistor.fm/episodes/tales-from-...
[1] https://oxide-and-friends.transistor.fm/episodes/the-sidecar...
[2] https://oxide-and-friends.transistor.fm/episodes/bringup-lab...
I'd hope your target audience would understand the limitations of such a thing, and I'm probably not the only person who'd rather read than listen even with the obvious caveats.
(these days automated transcripts seem to be no harder to mentally fix up the errors in as I read them than "somebody typing fast on a software keyboard and suffering the inevitable tyop and autocorrupt related issues" is, though of course others' mileage may vary)
It also seems like this was a spontaneous initial conversation and not part of the process, so I’m not sure why you are suggesting that they made it up.
But on the other hand, Oxide is very up-front about it, and their CTO is happy to go on HN and chat about it. So not knowing about it makes you look like you didn't do your research before the interview, or knowing about it and trying to force the issue anyways makes you look kind of arrogant (if you don't agree with it, you can just not apply).
Yes, it was an initial conversation as I said. When I asked about equity compensation I was told that compensation discussions are to be avoided because bringing it up could be considered a negative by the company.
I was using “interview” as a generic term for applying to a company, not literally referring to your internal process.
I didn’t interview with or even pursue Oxide after the conversation (and never said I did)
This all came up because I asked the person what the equity compensation was like. Not an unreasonable question when talking about a startup. That’s when they started advising me that I shouldn’t bring it up and it’s not something they talk about. After that I got uncomfortable about pursuing a company that discourages any conversation about compensation to the extent that someone felt necessary to warn me about it when I hadn’t even applied.
> no one at Oxide would tell you to "not bring it up" because everyone at Oxide knows that it is a subject dealt with early in the process.
They were trying to tell me that it was important that I avoid giving the impression that I cared about compensation, as that would be a negative if I talked to anyone else at the company. Just repeating what I was told.
> And finally, this is all assuming that you were talking to someone before March 2021, when we published our blog post on it.[1]
No, they told me to look up the blog post, but I had not read the company blog before taking to this person.
> After the blog post, compensation simply doesn't come up: everyone has seen it -- and indeed, our approach to compensation is part of what attracted them to the company!
Or maybe your approach to compensation is what filters people out of the application pipeline? I don’t think it’s realistic to think that this compensation strategy is what attracts people to the company rather than pre-selecting people out.
I read the blog post, but I feel like I’m missing the equity portion of the conversation still.
Regardless, is it so hard to believe that compensation “simply doesn’t come up” because potential candidates (like me) are sometimes coached to not bring it up? Or that the company’s stance appears to discourage bringing it up? This feels like some circular logic: Nobody brings it up because we discourage people from bringing it up.
I think you are missing that there are other reasons to avoid Oxide that doesn't have to do with optimizing for compensation. They may very well be optimizing for something else, but compensation is still a data point and while they may take less money, maybe not too much less.
Proofreading would set up an expectation on the part of readers that it -had- been proofread and corrected and therefore a commitment to perform a repeated "boring but important" task going forwards for whoever's doing said proofreading.
That way would likely lie either delayed transcripts or never getting to initial activation energy to provide anything at all.
So I think "add a quick bit of code to your podcast publishing workflow and a CAVEAT IN BIG LETTERS" is better to do first.
If it turns out enough people care about the transcript, doing it a more labour intensive nicer way later is something they can decide, well, later.
Shipping is feature zero, as ever.
This came as something of a surprise to me - six months ago I'd likely have been enthusiastically agreeing with you.
As an example, the transcript tab on https://www.thebulwark.com/podcast-episode/tom-nichols-jack-... was pretty readable to me in spite of the errors. Whether you'll find it the same is, of course, a separate question.