Google Wins Appeals Court Approval of Book-Scanning Project(bloomberg.com) |
Google Wins Appeals Court Approval of Book-Scanning Project(bloomberg.com) |
The typical librarian actually doesn't fight for the access of materials but fights to enforce copy right stricter than the copy right law even calls for. At my college I took down the sign that said "No Copying of Books" at the photo copier and I put in place the actual copy right law. I can't tell you how many visiting librarians were "shocked" I did that. Well we also allowed bottles of water and talking in all areas except the study rooms and areas that had doors, but that is another story.
The Law of Fair Use: The ultimate goal of copyright is to expand
public knowledge and understanding ... while authors are undoubtedly
important intended beneficiaries of copyright, the ultimate,
primary intended beneficiary is the public ...
[1] https://assets.documentcloud.org/documents/2461545/agvgoogle... [PDF]I'm not so sure I won't be aging myself by saying this, but I feel very nostalgic about libraries, and the thought of them becoming obsolete saddens me, not so much because I can't imagine a future with mostly digital books, but because I always felt at home at the library (and not in the "there's a man sleeping under the 600s shelf" sort of way). I always felt a sort of wonderment at the library, and that others were there for the same reason added to that. I didn't visit the public library for school reasons, typically. Instead, I just sort of explored. I have always been a slow reader, so finding the exact right thing to read next was a bigger deal for me than just reading book after book.
I also volunteered as library aid in 12th grade, which was fun.
Are there others that feel / felt the same about libraries?
DRM, copyright and the cultural myth perpetrated by big media producers that every tiny bit of content must be paid for are real the threats to libraries.
Knowing all the different information that could be hidden inside these walls of books that I was surrounded by that might be interesting to me was also something special. You can get that with the internet but there's really no visual, material feeling you get actually seeing the amount of things you can learn about.
Once/if you have kids, you'll get to relive that joy of libraries all over again. :)
(unrelated: my school libraries also prevented excessive photocopying of books.)
If you asked them why, they were decent enough to provide some sort of reasoned argument based on the actual copyright law to explain this policy, so I don't think they were just being jerks.
[This was in the UK, so I dunno if the laws were worse or better than in the U.S.]
Second, I don't think traditional librarians know how quickly the end is coming. Libraries will still exist of course, possibly as makerspaces, possibly as community centers, but with collections like "Library Genesis" [+] and storage continually to plummet in price, in the next 5 years entire libraries could be carried in your pocket.
EDIT: I'm not saying librarians aren't necessary! Quite the contrary! I believe their roles are going to shift to be advisors and guides. I'm saying the idea of the library as a place to go get knowledge itself is going to tail off, since its available over the Internet Firehose.
[+] "Library Genesis is an online repository with over a million of user-contributed books and is the first project in history to offer everyone on the Internet free download of its entire book collection (as of this writing, about 15 Tb of data), together with the all metadata and code for webpages. The most popular earlier repositories, such as Gigapedia (later Library.nu), handled their upload and maintenance costs by selling advertising space to the pornographic and gambling industries. Legal action was initiated against them, and they were closed. News of the termination of Gigapedia/Library.nu strongly resonated in academic and book lovers’ circles and was even noted in the mainstream Internet media, just like other major world events. The decision by Library Genesis to share its resources has resulted in a network of identical sites (so-called mirrors) through the development of an entire range of Net services of metadata exchange and catalog maintenance, thus ensuring an exceptionally resistant survival architecture."
We've got the resources now to put libraries in our pocket. Phones, tablets, etc can do so easily. At the same time, a library is more than the sum of its books, and needs more than just metadata to properly operate.
Yes we have a fire hose of content and people need people like librarians to help utilize that fire hose more.
When I left (I quit the job) being a librarian in 2008 more than 50% of librarians had lost their jobs in the next 2 years. Most people who handle budgets have your same idea.
My school district of 19,000 students had ONE librarian for all the elementary schools. CRAZY
Every time any of this is spoken it is in terms of job security. If the place you are employed at is sued you could then easily lose your job. It happens not a lot but enough to scare every librarian. When I proposed we leave the $4 a book library management system to a Open Source Evergreen the words spoken behind the companies was hysterically wrong but enough to scare anyone from using them. Back then Open Source meant evil and bad to librarians due to companies selling them goods spreading FUD.
Authors Guild v. Google Inc., 13-4829, U.S. Court of Appeals for the Second Circuit (Manhattan)
The ruling applies, of course, only to the United States, and only if it is not reversed by the United States Supreme Court. But I think the argument that Google's use of the book content is not market-destroying for book authors is correct, as I have bought books after discovering them through Google Books searches.
As a new academic librarian and recent graduate of UIUC's iSchool (top library school in the country[0] and an active hub of computer science), my experience is that the strict and shushing librarian is, mostly, as you all are recalling, a thing of the past. Libraries are more and more about collaboration, computing, and guidance. Kitchens more than grocery stores. Librarians are more and more facilitators, teachers, and collaborators. The open access movement has a very strong base in libraries.
[0] http://grad-schools.usnews.rankingsandreviews.com/best-gradu...
I believe that this would also clear up a service that scans your books and sends you the digitized version. That would seem to be fair use of my own library as well. I spent about $1200 getting roughly a 1/3 of the volumes I've collected over the years digitized at 600 DPI so that I could have them all available on my iPad for reference.
I've also acquired a nice guillotine paper cutter and a Fujitsu ScanSnap 1500 and have probably scanned 40 or 50 "trade" paperbacks with it, 6 years of Scientific American, several years of Air & Space, Nature: Materials, and assorted other magazines.
Would it be okay to have 30 seconds of any song available free online, just so that you could search for a lyrics of the song you heard on the radio, and then match it with 30 sec audio clip to find out whether the song is the one you were looking for.
https://www.techdirt.com/articles/20151016/08010632559/appea...
http://arstechnica.com/tech-policy/2015/10/appeals-court-rul...
And the Court's actual written opinion:
http://www.ca2.uscourts.gov/decisions/isysquery/c3458e0a-f3d...
Note, though, that Google in the middle years of this dispute sought to acquiesce to a class-action settlement with the Author's Guild. That would have more-or-less abandoned the (strong and ultimately successful) fair-use argument, and set up a system where Google and the Author's Guild were economically aligned, with a precedent against other (less deep-pocketed) groups who might want to make a similar fair-use argument in the future.
Third parties including the American Libraries Association, EFF, and ACLU objected to the potential negative effects on competition, privacy, and free-speech of that proposed settlement, which helped prevent it from being accepted by the courts. That forced Google to fall back to its original defense, and led to this broader win for fair-use principles. For more details, see:
https://en.wikipedia.org/wiki/Google_Book_Search_Settlement_...
Although Google says that they're using it for one specific purpose, I can imagine that they'll use it for other things like improving their search and ad technologies. If that's the case, then wouldn't Google's work be considered derivative of the original content?
Why should the authors involved not be able to re-sell those digitalized versions to other companies (especially other search engines)? There are probably lots of companies that would like that data set and be willing to pay for the use of it (including me).
The ruling pretty much addressed this already:
"Plaintiffs’ contention that Google has usurped their opportunity to access paid and unpaid licensing markets for substantially the same functions that Google provides fails, in part because the licensing markets in fact involve very different functions than those that Google provides, and in part because an author’s derivative rights do not include an exclusive right to supply information (of the sort provided by Google) about her works."
If that's the opinion concerning providing search in the books, I think it's highly likely that the same logic would apply to improving search and even ads.
[1] https://www.techdirt.com/articles/20140114/10565225874/copyr...
If true, I'd count this decision as a net negative. (Sorry Google, your balance sheet means nothing to me.)
I'm interested what kind of evidence played a role here. How does the judge determine it didn't harm authors, other than using a time machine?
For instance, if I have a bunch of books about drawing, I'd like to scan them all so that I later group all of the figure drawing pages in one folder, all the gesture drawing pages on another, etc, so they can be more easily used (and more useful) as reference.
Does anyone here recommend a way to scan books at home? I'm not against buying a contraption.
If yes, then ever since copyright law was enacted. It's called fair use and the public domain. Also note not all artistic creation is actually copyrightable.
It doesn't. However, this isn't about rights shifting from the owner of a copyright interests to someone else, its about power to exclude particular uses that the owner of a copyright interest never had (note that statutory fair use is largely a codification of pre-existing case law on fair use which was grounded in the First Amendment, and so generally addresses powers that not only the copyright owner never had under the law, but powers which Congress could not give the copyright holder, because doing so would violate an explicit Constitutional limitation on the powers of Congress.)
I am blown away how this is ruled legal, but in many cases scanning books as an individual is considered copyright infringement. As though they don't reap financial benefit from expanding the scope of the data they control? It's their entire business model...
At least in the US, you are overstating the certainty. While I agree with you that it should not be infringement, prior to this ruling, the legal status of scanning books you own for personal use was not clear.
The qualified expert opinion I got regarding personal scanning was from a law school professor specializing in copyright (and who actually participated this case), and was along the lines of "it's probably legal but there is not yet settled case law".
This ruling probably moves it closer to settled, but I'd be suggest against placing large bets unless your qualifications exceed hers.
Practically, the result of rejecting the settlement seems to have been a slowdown in the pace of digitization, and that readers are left with still no way to easily access orphaned works.
Anyone with shallower-pockets that then tried to do what Google did would likely have been sued by Authors' Guild – now strengthened by Google cash – or other members of the class. There was no precedent or requirement that others be offered the same deal as Google: if Authors' Guild liked their deal with Google (and why wouldn't they), they could tell others, sorry, we've already got a system in place, you're not part of it.
But further, why should other 'little guys' have had to fight a legal battle with Authors' Guild, or negotiate under threat of litigation by a de facto Authors' Guild-Google alliance, just to do something that (now, finally) is clearly authorized by fair-use?
The class settlement's Google-financed-and-managed system, for the benefit of the Authors' Guild class, would have started with an overwhelming and likely legally and economically insurmountable advantage in the scanning and marketing of older books. That gave rise to the centralization and privacy/censorship concerns of the ACLU, EFF, and American Libraries Association. They're smart and like old books, too – but perceived a risk that outweighed the benefit of "just scan 'em all quickly – under a Google/Authors' Guild monopoly".
https://www.youtube.com/watch?v=4JuoOaL11bw https://www.youtube.com/watch?v=3lLL0mUZHwU https://code.google.com/p/linear-book-scanner/ http://linearbookscanner.org/
If I were to do more and had the space, I would have gone the diybookscanner.org route to improve the quality and processing rate. At one point I belonged to a hackerspace in Oakland that had one available.
Post-processing workflow was much easier, and involved using scantailor (awesome free software to batch align, crop, white balance, etc the pages) and then Acrobat for OCR.
Emeritus community hero Daniel Reetz spent 6 years creating the "Archivist" scanner [1]. He and his collaborators have done a phenomenal job, and created some of the best documentation I've seen for any project (open-source or otherwise). The "Lessons Learned" front matter alone is inspiring [2].
So far I've found that book scanning is an ideal "DIY" project: enough hardware & software quirks that are gratifying to puzzle through, but nothing super difficult. In fact, it is exactly like building and calibrating a simple scientific instrument and learning to collect and process image data. To @planfaster or anyone who is considering book scanning for private use, definitely do it!
I highly recommend buying the "Archivist" scanner kit + electronics pack available at http://tenrec.builders/. There is ample hard-earned wisdom in the forums and tenrec supplemental docs about dozens of minor process details where you think "Why don't people just do X?" and it turns out X isn't ideal, and neither is Y, but Z works fine.
The main thing that I didn't consider before starting was that the scanner hardware only facilitates one very specific part of the workflow: taking pictures of flattened pages with (nearly) identical resolution and positioning. It's an important step, and reducing it to 5 seconds per page doesn't magically eliminate tedious downstream processing with other tools[3]. All that said, it's very rewarding, and really fun to start thinking about what you can do with scans, e.g. turn entire books into posters [4].
[0] https://news.ycombinator.com/item?id=10070529
[1] http://www.wired.com/2009/12/diy-book-scanner/
You know, I might have agreed with you before. But after watching the Vanity Fair interview with Elon Musk and Sam Altman, and Sam talking about how no one person really understands how Google's first page results are created anymore, or how machine learning is matching people on dating sites (and those people are having babies, determined by machine learning), I think we're approaching an inflection point where algorithms will be able to provide that guidance.
We're not there yet, I agree with that. But we're closing fast on that future.
Why wouldn't the Authors' Guild be willing to offer others the same terms? How would it benefit them to depend on Google? The agreement would have granted a new sort of status to Google, but there is no reason that status would have to remain unique.
Regarding the anti-trust angle, the wiki article cites an MIT paper that concludes the settlement would not violate anti-trust and would in fact generate a consumer surplus. [1]
Forcing Google to fight for fair-use may have been a sound Machiavellian strategy for the EFF (as it resulted in today's ruling), but it's ironic then that the main reason why the settlement was rejected seems to be because it was not strong enough on copyright, as expressed by individual authors' concerns over loss of control, and freeing of orphan works.
[1]: http://www.criterioneconomics.com/Google%20and%20the%20Prope...
Given the explosion of niche interests over the past 50 years I don't think its feasible for librarians to provide that service except at extremes of the spectrum: internal university or corporate information of extreme specificity and controlled scope at one end, and public topics of general specificity at the other.
In the middle is a vast information space traditionally only explored and indexed by clubs and societies but now predominantly by the search engines.
Think about asking your local public librarian about functional programming, or the history of turbine-powered cars. They'll have to refer to a index of recommended texts[0] and hope there's something vaguely similar there, which is far inferior to a search engine which can access vast reservoirs of specialist discussion on such specific topics.
[0] I can't remember the name of this index after all this time but it lists 'go-to' references for a long list of topics.
Would you mind explaining what you mean here in a little more detail. I'm curious to learn more about this.
Or we need to get better at automating the process of helping people find what they're looking for. That is, keep improving search engines.
More likely, it was a rule adopted because fair use analysis is generally helped by using a limited portion of the copyrighted work, and organizations concerned with liability don't really want everyone independently trying to figure out how limited a portion is limited, so they like to set some standard that is likely to be limited enough in most real cases of the type they are likely to be exposed to (e.g., nonprofitable educational uses, for school projects) as to mitigate risk sufficiently.
> Meanwhile we just had to cite where entire photos came from! I think it was just the recency of the music piracy issues that made them care.
That's actually perfectly sensible -- the demonstrated propensity of interested parties to file a lawsuit, and the likely damages in the case a suit is lost, are perfectly rational factors to consider when determining how to craft a legal risk mitigation policy.
Any book you cannot scan?
Unfortunately, WE CANNOT SCAN ANY PUBLICATION BY McGraw
Hill. They do not allow us to scan their publications. When
we receive any publication by McGraw Hill from customers,
we will return it back by charging actual postage fee."
http://1dollarscan.com/faq.php#8aFor a textbook style book "chopping it in half and throwing it in a scanner" is a bit more work than the sentence would suggest. The most cost effective scanner for this is the Scansnap 1500 as it will scan both sides of a page, has a 100 sheet "feeder", and will OCR the text (using ABBYY which is included). It screws up occasionally and especially on magazines which are very thin / shiny paper it can take a while (and several rescans) to get the magazine scanned. So in general there is a pretty solid time advantage to using 1dollarscan. Especially if you can use your nights and weekends productively doing something else.
That said, once I didn't have another stack of 10,000 pages to go at the end of the month (I had scanned all the "obvious" targets, minus the McGraw-Hill books which they won't scan) I did switch over to manual mode with my scanner because while the total cost of the cutter and scanner was close to $2,000 (not quite 2 years worth of 1dollar scan services) it is a capability that can sit idle without too much cost.