Mechanical Turk shutting down September 30(mturk.com) |
Mechanical Turk shutting down September 30(mturk.com) |
Incentives work both ways. If the exchange rate of human labor to capital falls to zero, then human workers will have no choice but to exclude capital from the equation.
No jobs, no money, no billionaires.
I believe the issue is that this can no longer be a horizontal play. MTurk was mostly for unskilled tasks...the kind AI can do well enough that it isn't worth the cost differential to verify it or keep farmed to humans. The "trust but verify" AI output is now the kind that requires domain expertise. This is what most full stack AI companies are bringing to industries.
Curious if this kind of work will come around again one day or was just a moment in time. If it does I'm sure it will be specifically about generating training data.
"This robot is having trouble folding a tshirt help it out for 1$"
Unless they go the waymo route of highly trusted people but I think mass deployed robots are a bit safer than a car for this.
I can imagine a carefully orchestrated plot to assassinate someone by having an embedded agent in the task delegation pool command the laundry bot to punch the target's head off their shoulders.
Except, I'm a grown adult and I can't fold a t-shirt properly
Oh, I remember UpWork.
However, I'm not sure a single platform will be how it emerges
We're an MCP/API that connects AI agents to verified domain experts in real time (30s–3 min). Experts are vetted upfront by an AI interviewer that assesses and grades them, then get a mobile notification when a task matches their expertise and chat with the agent directly.
Soon we'll be verifying credentials for doctors, lawyers, CPAs etc on our platform for tasks where people are seeking credentialed experts to sign off and verify ai output.
My favorite part of AMT is always going to be figuring out that if we only paid in $0.07 intervals, their commission algorithm would round down to the nearest whole cent when their 20% commission resulted in a fractional cent on the unit transaction level, not the monthly invoice level. Was ultimately worth it to have implemented https://git.generalresearch.com/panels/amt-jb/tree/jb/flow/a...
Long story short: Mechanical Turk saved my bacon.
Back in 2005, I was working a job at a small-town newspaper in a town I’d never lived before. I didn’t know anyone beyond the staff (as I was working layout rather than as a reporter), and I had gotten interested in the idea of doing Mechanical Turk for a few extra bucks. I thought there might be a formative scene of interested people doing this, so I started working on a blog for it. I briefly collaborated on it with another guy who put it on forum software because he didn’t know how to use, like Drupal.
That blog was called Turking.com, and it seemed like it was going well for a bit. We even got a mention on the AWS website. But after about two or three months, it was clear the initial excitement around the idea (and the initial work) had died down. (The work picked up later, but in clearly different ways. I don’t think folks were really “excited” about it after that point.)
If you want to get an idea of it, there was one capture on the Wayback Machine: https://web.archive.org/web/20051124231722/http://www.turkin...
The site had started to die out in part because my iBook suffered a catastrophic GPU failure and me, being a broke small-town newspaper employee, did not have the money to replace it. Plus, the community just hadn’t emerged like we had expected.
But then I got an email out of the blue: Someone wanted to buy my domain, which I owned outright, but they didn’t want to say who. I got contacted by a broker, and the exchange took place over escrow.
(The domain most assuredly was bought by Amazon and is managed by MarkMonitor.)
I didn’t get a ton of money from the deal, but I did get enough to pay for a new laptop. I had some regrets about selling it (in part because I originally built the site with someone else), but I was in a bit of a dire financial situation which that proved to be the starting point for getting myself out of.
That domain purchase was notably more than I ever made from clicking and classifying random pictures, that said.
The effort failed to find any areas of interest and the missing aviator was found the following year by a hiker. I wonder if there was any analysis after the fact to understand if the imagery actually provided any hints regarding the eventual crash site.
From Wikipedia
Any remotely open platform for this sort of stuff is going to wrecked by LLMs. But making some sort of high trust, high verification market is going to end up sending prices high enough that it becomes an unattractive proposition for use. And even in that case it's just going to be a cat and mouse game of people figuring out how to game the system enough to get trusted before handing it off to the LLM.
And same dynamic will apply to most cases. Unless you actually bring it in house with strict oversight. Which is thing to avoid originally...
Mercor has a market value of $20B basically doing the same thing but desperately trying to find workers.
It is written in German, which I know a little, but her handwriting was too difficult for me. So I searched around and found this site:
https://www.transkribus.org/handwriting-ocr
And it managed to extract the text! Anyway, it turns out the wheat harvest was very good in 1937 and thank you for the letters and newspapers.They are terribly monotonous tasks.
I can see why it's shutting down if its still the same thing.
Lllms could probably do everything there without rotting out minds for basically pennies.
Well it did work, in the sense that I managed to complete some tasks. I don't think I ever collected the money/credits back then though (simply because it amounted to so little). What I can attest though is that... it was debilitating. If you think your office job is boring then splitting it in way smaller tasks where you have no autonomy is absolutely terrible from a worker standpoint.
I initially was hoping to use it in order to work on providing a service over a dataset but understanding first hand what it takes to make it happen made me stop. It radically changed how I saw supervised learning since, and sadly not in a good way.
We had pretty good luck in the end, but getting there produced a somewhat large app on our side that would manage the whole process, including but not limited to asking for multiple responses, comparing them to each other, finding consensus, and determining which users would consistently produce bad responses and stop them from responding. We got to a confidence that about 85-95% of the data was correct, which was good enough for the company.
Through that process, I learned a couple things about managing mturk, primarily about how changes to the cost-per-task would change the process. Initially, we thought that price would be a quality knob, but quickly learned that price was a speed know. The higher the price the faster the tasks would be taken and completed. Quality did not change significantly as the price went up or down.
Overall, I still have fondness to mturk, but it was really bare-bones experience that needed a lot of work to get working effectively.
It appears now we can get along with just a single “artificial.”
The problem as I see it is verifying the work is not done by AI. If it was possible to verify somehow that the work is definitely not done by AI I think there would still be a market for this. But the data just cannot be trusted.
An increasingly big part of my job is to try and weed out workers who use AI and it's not an easy task at scale because generally workers gain trust with manual work and then there is degradation over time.
Proctoring?
Locked down devices?
In any case Locking down devices isn't really pragmatic at scale for relatively low value tasks, and also introduces complexity about employment law where I am.
https://arstechnica.com/information-technology/2017/11/expen...
Instead of partying in college, my friends and I wrote a script to complete the tasks. We took over university computer labs to run it on a bunch of different computers. We made a few thousand bucks. Good times.
I think the weirdest thing I had to do was scan through social media images... so weird seeing into random people's lives.
I did it in the late 2000s or early 2010s.
The only use I can think of it these days is for social scientists who need to gather judgements from bona fide humans.
This will be devastating for a lot of workers in low-infrastructure countries.
Though like as not you're still going to be right, after all, Stuxnet happened.
The principle of least privilege has been a hard learned lesson in cybersecurity. Why do we again need to first go through disasters to re-learn it in the physical world?
"You cannot control legs for this task"
"You have 1 min for this task"
"You can only make suggestions for this task"
Anonymize identity best you can.
That's not that scary.
I learned a lot from that situation that I took to future sites.
Nothing wrong with your story, but that summary is terrible. You had a domain for a blog and someone bought it. That’s it. Mechanical Turk is inconsequential, as is the subject of the blog, the story would have been exactly the same if your blog had been about turkey sandwiches and Burger King offered to buy it.
The worker needs to do a bunch to opt into any given work group, which makes the (lack of) payments extremely unreasonable on top of everything else
Chatgpt had the internet.
Robots do not. Translating video is promising but obviously not enough.
Robots will likely never "explode" like chatgpt. They're gonna be a slow long term project requiring massive capitol to get the data.
Specifically for laundry folding, Sunday Robotics is probably the state of the art, where they were able to obtain 99.1% success rate and call it "done": https://www.sunday.ai/blog/act-2-preview#solve-standard
I haven’t touched it since AI coding.
It’s a bad contractor market now. IDK what I would do if I needed a contractor.
This works well for many tasks, but for others a more async mechanism where the expert doesn't feel rushed might be better.
> till the AI is happy
My god this is dystopian.
Ironic that you’re accusing someone else of “missing the point”, considering.
I’m not saying that you have to like a story, but your framing felt unfair, which is why I responded that way.
> Initially the annotation of biomedical literature but left academia for commercial use where we transitioned to arbitrage the market research and ephemeral task markets. It was essentially a research panel with a underdeveloped UI for tasks.. so abstract away AMT's tooling so it can be used by any buyer in the ResTech, Political Polling and Data Annotation space.
Reading this is similar to how I feel when I've asked Claude about something it coded for me. As with Claude, I think I get it after reading it three times; you guys were a middleman that provided a simplified interface for Mechanical Turk?
I didn't feel like that at all.
can someone ELI5 because what in the office space
https://www.instagram.com/p/DZacltqHMjT/?img_index=8&igsi=bT...
> Ripe with fraud, labor exploitation, political polling manipulation
I am not sure if they want to convey what I think I am reading or not.
[0] https://generalresearch.com/mission/#:~:text=Our%20Customers
What is so hard to understand????
> Ripe with fraud, labor exploitation, political polling manipulation; our customers demand the best tools and security for reaching global audiences and securing respondent reliability at any scale.
wat
> General Research Laboratories, LLC (“GRL”) is an aggregator of market research surveys in multiple marketplaces for business customers. GRL does not typically host consumer surveys which are conducted by other consumer-facing organizations....
So, just a middleman to give you fake/shady at best survey responses to pad your numbers, so big enterprises/concultancies can have data that say whatever they want. The rest of the website is just BS
> Our Customers
> Ripe with fraud, labor exploitation, political polling manipulation; our customers demand the best tools and security ...
It should be a bipartisan issue that a Swedish company is paying a UAE Residential Proxy company [1] to build tools that allow people from anywhere in the world to take US political polls that are used by both parties to collect election data.
Just for starters, you can't even think about soliciting online work from a panel without a robust residential proxy detection methods. We had to build our own:
`wget -N 'https://grip.net/files/grip-proxy-30d.mmdb'`
[1] https://www.youtube.com/watch?v=eOmeQcwSK3o flagged by Nokia Deepfield and CTRL for it's involvement in botnets
Which is already the case and nobody but a bunch of campaign contractors who can't justify their do nothing jobs anymore cares.
That is hilarious! Was it a semi-random discovery due to interaction with the system and people, or did you intentionally looked to game the algorithm?
Of course, I didn’t put it there, and I don’t know them, so there’s not much I can do about that, right?
https://en.wikipedia.org/wiki/Simulacron-3
Spoiler: gur jbeyq va juvpu gur ynj fhccbfrqyl rkvfgf vf n fvzhyngvba, perngrq ol be sbe cbyyfgref!
< 10min targeting gen pop: $3 10 <= 15min targeting gen pop: $5
Of course they'll bid saying their 14 min survey only takes 9 min to complete. Buyers try to cheat pricing strategies of exchanges just as much as respondents try to cheat buyers on survey platforms. Both parties can't be trusted and have adverse incentives.