Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Public money does go to the costs of that.

That’s where most of the PDFs come from.



Hmm, no. In the case of old print journal archives scanned by JSTOR, it's JSTOR paying the cost of the scanning.


Anyone can scan, host, and curate - including one or many peers.

We do not need any entity in between researchers and other researchers.


Practically speaking, no-one is going to scan centuries' worth of historical journal articles for free. And JSTOR isn’t really an entity outside the research community; it’s a non-profit that grew from within it.


I think we’re talking past each other mostly because I start to lose value on most research for my topics of interest that are more than two or three decades old, let alone five or six.

This may be different for others.


So, this discussion started with me saying that:

> [JSTOR is] a non-profit that scans old journal articles that would otherwise be a huge pain to access

You'd think that would pretty obviously include the vast swathes of material published before the internet and PDF preprints were a thing (and the considerable amount of subsequent research that just never got uploaded to anyone's website, for whatever reason).


Okay, so why can’t the volunteers scanning for Anna’s Archive perform this again?

Why does it need to be technically centralized and politically weak?


Anna's Archive is primarily a meta search engine. I see that they've put out a call for volunteers, but I don't think any significant part of their archive consists of volunteer scans.

As to your 'why' question, we're talking about boring, thankless work that in many cases breaks the law. It's not exactly surprising that we don't see millions of people signing up to do it for free!

I don't understand your second paragraph.


I simply don’t see JSTOR outlasting (as an organization, or a technology/archive) individual contributors and torrent trackers.

Solution which does not require millions of people =]


At present I don’t see individual contributors making any significant contribution towards scanning old academic journals. Why would this change in the future? There was at least a decade where the technology to enable this existed, and where most older issues of most journals were not available online, and yet the army of volunteer scanners failed to materialize. Now that there is already a non-profit doing this archival work at very high quality and with economies of scale, it seems even less likely that such a volunteer effort will materialize. But if you want to prove me wrong, get off HN and go scan some old journal articles!


Probably just going to keep reading the news ones,

and if I want to synthesize information from old ones, I’ll query a robot.

This is still research we’re talking about, right? Not a first-edition of famous literature?


'Research' includes subjects such as history, where old documents are important for obvious reasons. Besides such cases, I gave a concrete example above of a paper from the 90s that's available online because JSTOR scanned it. Hardly ancient history.

>I’ll query a robot.

Which has read all the old papers that JSTOR has scanned. That's why it's able to synthesise them for you!


> I gave a concrete example above of a paper from the 90s that's available online because JSTOR scanned it. Hardly ancient history.

For certain sciences, 36 years is two lifetimes. Like I said: we’re talking past each other due to our research needs.


We're talking past each other because you only care about your own research needs and apparently can't see any value in opening access to documents which other researchers might need. I am not a historian, but even so, it's not lost on me that a historian might want to access some old documents. Just because you think that JSTOR isn't useful to you personally (though, I can guarantee that your 'robots' have been trained on it) doesn't mean that it's not serving a useful function to the academic research community as a whole.

One can imagine the quality of research in fields where researchers can't be bothered to read anything that wasn't published in the last five minutes, and rely on other people's partial and possibly inaccurate summaries of key results. But that is another topic.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: