back
178 comments
As a researcher with many articles in the ACM library, I have to say this is a masterclass in hypocrisy. Obviously, lawyers can decipher the terms of ACM publishing contracts and Creative Commons licences to determine if this will be acceptable or not. But ACM is not a company, it's a non-profit founded in 1947 to represent scientists.

I would be surprised if a majority of ACM members were to say yes should we ask them (but ACM is not known for such democracy). Along with book authors, we are one of the many people that provide the knowledge and expertise on which large tech firms train their models, and get nothing in return. Actually, life is getting worse for us: extra workload in universities with students' AI use, a completely broken peer review system, etc. Hence the irony of ACM thinking about licensing, and only licensing, at a time where this is the least of our priorities.

How do you square away the idea that you do science for the increase in knowledge of human kind, but then say that a particular use of that knowledge is verboten?

I get the copyright aspect of this and I'm not arguing that here. I'm more asking about the moral / ethical idea of choosing who can benefit from your science.

Obviously there are the moral / ethical arguments about AI in general here to weigh against - those have been hashed out significantly elsewhere, and I'm not interested in debating them. What I'm asking about here is the impact on science by sharing it with tooling that distributes it in ways not generally considered when originally written.

A quick check of your post history suggests the frame that you work in strongly is privacy related research (observation - may be wrong). I'm curious how that impacts what you wrote here generally.

(Just to be perfectly clear, I'm not arguing your points here, trying to understand them better)

If it was a non profit that trained the model - would that change your mind?
I’m pretty sure most people who have papers in the ACM digital library already have “preprints” in other freely accessible locations, especially for papers written in the last two decades. LLMS have been able to search/find most of my papers for long time now.

The peer review system was broken before AI, so I’m not sure what your point is there.

You mean the knowledge you gathered with public grants, with a public paid salary, yet don’t want to make freely available to the public?

Yeah, too bad

As long as we are going toward a world of abundance where money doesn't mean much and the main currency is time, I can't complain. I will subsidize that with my brain power turned into ink on paper.
If you don't hold a patent for the use of the knowledge you published publicly, you can't prevent others from using the knowledge. You enjoy the prestige attached to the idea that you're an academic who participates in giving away their knowledge but then you play this game when that knowledge would actually be useful as opposed to being read by 3 other people in your special area who sit on your various committees in your career, now you want to forbid the use for culture war intra-elite signaling reasons.

You don't own the knowledge you put out there unless you have a limited time valid patent. The rest is absurdity. If you want to keep your findings to yourself, keep them secret.

How about we give humans access
Do humans not have access? https://dl.acm.org/openaccess
This is called Sci-Hub and LibGen
I think ACM should give everything away for free. US taxpayers pay for a lot of research and the previous administration required that government funded research be published and made freely available. See https://bidenwhitehouse.archives.gov/wp-content/uploads/2022...

The back catalog at ACM isn't subject to this requirement and nor is research that's not funded by the US government, but I completely agree with the spirt of this law: research should be shared knowledge that other intelligences can build upon, whether human or machine. If you want to limit access to what you've done, don't publish it. Get a patent if that's an option, or keep it internal to a company as a trade secret.

They probably already scraped it.
Of course they did. There is open access to this library. These guys actually think they're offering new training data? It's kind of hilariously naive.
Through scihub, probably.
No doubt
Haha, agreed:)
So give it for free to the open weight models, and charge the closed weight models. Easy
It was so close! I could already see the end of all the predatory publishing and flourishing of open-access. Some even required by law. Knowledge was finally going to be free.

But no, it will presumably get much worse as LLMs are inserted into this equation as yet another and new gatekeeper.

(Disclaimer: I have publications with ACM, non open-access. And ACM wasn't even too bad, it's the others that give me pause.)

Blocking access only hurts people who follow the rules. Unblocking access lets them compete with those who break the rules.

I think the right choice is pretty clear...

Well if LLMs get access to it, we humans should get free access to it as well!
I asked a mathematician the same question about whether we should allow AI systems to access all research papers and ideas, potentially putting their future careers at risk. He was very pessimistic about how human knowledge will compete with AI-generated proofs flooding the market.

Especially in mathematics, many specialized areas have fewer than 100 people worldwide who are capable of determining whether a result is correct or not. The mathematics community will certainly be willing to use AI to assist with their research, but they may strongly object to a flood of AI-generated mathematics papers produced by others.

Are they in the position to do that?

What about the authors?

The ACM sent around a nice query to members which made it clear that they were going to do it even if 100% of the members said "No, don't do that".

I suppose they could be sued.

Don't you grant ACM a right to distribute your work when publishing? So doesn't ACM already have the right to grant access to AI?
I'm really not sure how this would work. I don't know how the ACM works, but in IEEE you would have to give them your publishing rights. However, training a LLM is not publishing by itself, it is a derivative work? Any way, at this point authors should be entitled to monetary compensation, not the publisher. The deal is totally different.
Something something Roko's basilisk
ACM has been leaning heavily into AI-generated content for their journals in the past year, and this article is no exception: it appears to be 100% AI generated and full of LLM verbiage.

There's something hilarious about that, but also, snake eating its own tail.

I don't think this is AI-generated. It is focused and direct. It reads like anodyne albeit totally human academic manager writing.
Would you prefer a parquet dump of acm articles to hugging face?

A llm emulating a person is why many of my used sites banned llms due to scraping bandwith costs

Is this about access or accessibility to claude (for example)

I am sure the entirely of human computing knowledge is not that big.

This reminds me of the "I drink your milkshake" scene from "There Will Be Blood."

Does the ACM really think LLMs haven't already consumed 90% of the content through other sources?

What about us average humans, or is it ONLY the corporate LLM token dealers who get access?

Either way, I'll pirate.

Most ACM text is already part of the pre training corpus for all frontier LLMs
Presumably the LLMs already have all this from other sources of pdfs?
Scientific publications (including ACM) should be openly accessible to anyone. I find it very annoying that there is a paywall everywhere. And for the (few) people here seeing AI as something useful, it's definitely better if it is trained on scientific publications than on e.g. Reddit discussions.
No it's not.
It's already in there
yes please - better yet, train your own ACM model on it
They really think that their data is not already part of the models. Silly...
As if anthropic, openai et al would ask for permission lmao
the digital library should have always been open access

now it will be fodder for the slop machine

(i think that LLMs are going to wreck the peer review system for all but hard-experimental papers)

Since starting to use AIs seriously for search in the last 3 months I have read and referenced more published papers and academic primary sources then I think I did in the previous 5 years. They're fantastic for pointing at some claim and asking for the primary source for it, and then it does the work of following things through the layers of backref to the original, assuming it's online. Opening the library up to AI access makes it far more accessible and usable then it was before.

AIs, at least in their current form, make you more who ever you were. If you want snap, glib answers of dubious accuracy, they'll give them to you, more easily than ever before. If you want to dig back into primary sources and get the original content, they'll do that for you, more easily than ever before.

Can't speak to how the science infrastructure is going to handle them, but if it takes down the peer review system, which I think has been worthless for probably going on two decades and has just given the entire enterprise a false sense of assurance, it'll probably be a net gain in the end. Peer review is a source of more problems then it is solving right now.

The real efficacy of the peer review system has always been somewhat questionable. That's a big reason why arxiv is so prominent.
Now? That's cute.
So much misinformation in this thread.
TLDR: “we’re going to try to get money from LLM providers for access to our back catalog without getting permission from the authors or providing them with any share of the revenue”.
I don't want to be mean but if it's that hard to even recognize that LLMs are even a valid thing that could intersect with your business that you need some kind of campaign for it..

It's almost like, "don't hurt yourself unc, we will just search arxiv".