back

by artninja1988·2y ago·view on hn ↗
There is currently a project going on to create the pile v2 which has only permissively licensed data, because of all the bickering about copyright.
3 comments
So a bunch of extra work to create a downgrade? I'm sure that's going to be very popular.
The training data distribution is the only thing that matters, not the actual content
Unless you want something like style from a range of authors, knowledge of a fictional universe or storyline, or other domain specific data or style characteristics.

A blanket removal of copyrighted data would make a bot sterile, boring, unrelatable, and ignorant of culture and common memes. We have amazing AI technology. Let's lean into it and see where it goes.

I agree, hypothetically, if I were ever to have an opinion on the matter. Just playing to the audience and potential audiences if this ever gets read into evidence

;)

Haha... just kidding... unless.. ?

By the violating the copyright hundreds of authors.
if the pile contains the code to go from step 1 to step 2 and then to 3 then couldn't you just remove the parts you don't want from the raw dataset and re-run the code?
> because authors prefer to be paid for their labor

FTFY.

I asked an AI tool to create a cheery poem about ringworms infecting kids from the 1600s and it created something that's never existed before. Which author gets paid for this labor they performed?
Is it labor if it automatic? Which worker gets paid when I turn on my oven to heat my food?

I can see value, but no labor.

Naturally, but I wonder what writers are going to do when the AI trained purely on suitably licensed content is still good enough to make most redundant.

(The authors in best seller's lists may well be immune for a bit longer than others writers, as they're necessarily the top 0.1% of writers, but not forever: nay-sayers claimed that AI could never beat humans at chess or go because the games required special human insight).

Once upon a time, nay-sayers said that nobody could travel to the moon, regardless of what vehicle they used. They were wrong. Once upon a time, nay-sayers said that nobody could transmute lead into gold using alchemical equipment. They were right.

Nay-sayers who said that no possible algorithm could beat humans at chess and go? They were wrong. Nay-sayers who say that these algorithms cannot write better books than humans? Well…

> Once upon a time, nay-sayers said that nobody could transmute lead into gold using alchemical equipment. They were right.

Now I'm wondering if, with modern knowledge, you could build a 0.5 MeV heavy ion accelerator with only the things available to a medieval alchemist.

I'm thinking probably yes? Triboelectics can get the right voltage. But how good does the vacuum need to be?

> Nay-sayers who say that these algorithms cannot write better books than humans?

They may be right or wrong in the specific, but I think they're asking the wrong question, too specific.

By "these algorithms", do you mean the ones that currently exist, or the ones that will exist next month, next year, or in 2034?
We're not developing new algorithms all that quickly. My point is that one shouldn't dismiss criticism out-of-hand, just because some critics of some other thing turned out to be wrong: for this point to be valid, I don't need to be making criticism. On an unrelated note…

Personally, I'd be referring to the family of algorithms that purely take as input a context window and provide as output a prediction of the next token likelihood. (Plus or minus iteration, to generate strings of text.) Pejoratively, one might call these "fancy Markov chains", though as with most pejoratives, that's overly reductive.

All the approaches we're seeing marketed heavily are just fancy Markov chains. I expect every "new algorithm" for the next 5 years at least to be a fancy Markov chain, because that's what I expect to get funding. (I do expect that some people will be working on other approaches, but only for amateurish reasons.)

These are fancy Markov chains in the sense that humans are just chemicals and computers just do math. Technically true, but not even "overly reductive"; it is just wrong if it is used to imply that, e.g., humans just swirl around in beakers or the most complex thing you can do with computers is trigonometry.

You can make anything sound unimpressive if you describe it sufficiently poorly.

And: So many different variations are published every month. There are a good number of people in serious research trying approaches that don't use cross entropy loss (ie, strictly next-token prediction).

I don't know what the trajectory of the technology is over the next ten years, but I am positive no one else does either and anyone who thinks they do is wrong.

Strictly applying the definition, the entire universe is a Markov chain (thanks to quantum discretization!) People who use "Markov chain" as a pejorative are just idiots.
Dunno. Writing fiction myself; asked AI to read it aloud. Narrative paragraphs worked fine: a clear, if a bit deadpan, slightly tone-deaf delivery. But dialogue was horrendous: it didn't understand emotional reactions and connotations at all. More so than cringey and robotic, it felt soulless. And the distance from "something that makes sense" to "something that feels human" felt unsormountable. Yes. Many novels will be written with LLMs in the coming years. They might even touch us. But this little Text-to-Speech experiment felt like an evidence that this technology has a void at its core: it doesn't have access, like a human does, to a gargantuan emotional spectrum, which allows us to understand all sorts of subtleties between what is being said, and why, and what does it actually mean, and why does it affect us (or, hell, how should the next line be read in this context, because it has no context, it doesn't feel).
I'm also writing a novel, and using text to speech to hear how it sounds. One of the ones built into Mac OS. And I'd agree with your assessment, I value the synthesiser for bringing my attention to things my eyes gloss over, such as unnecessary repetition and typos which are still correctly spelled words (a common one for me is lose/loose).

But: AI was seen as "decades" away from beating humans at go, even 6 months before it did.

I don't know how far we are from them writing award winning novels (awards we care about, it doesn't count if it's an award for best AI), though my gut feeling is we need another breakthrough as significant as the transformer model… but even then, that's only a 1σ feeling.

You wrote a book with an LLM and it has bad dialogue _today_. So why do you act as if all this progress hasn't only been in the last couple years? Fast forward 10 years from now and you think LLMs won't have the capability to write compelling fiction, that what we have then will be exactly what we have now?
> The authors in best seller's lists may well be immune for a bit longer than others writers, as they're necessarily the top 0.1% of writers

The top best selling. Only one of many possible reasons for that might be the quality.

Quality is subjective, therefore I think it is reasonable to say the best are those most able to profit rather than, e.g. winners of the Nobel Prize in Literature, or the list of books people most pretend to have read.
I disagree. While quality is subjective popularity is different to quality. Otherwise we would only have marvel movies.

Your argument assumes no marketing is manipulating the best selling lists, something known to happen in new york best selling and others.

Regarding people pretending to read "fancy books" (my term) I think most people just dont read, but I find it annoying, for ex. when people see my shelves, that some people think that I buy those books to impress somebody. It is as if people that do not enjoy reading cannot conceive somebody enjoying it. I think it is slightly antiintellectual, and cynical. I have a better opinion of people in general.

> I disagree. While quality is subjective popularity is different to quality. Otherwise we would only have marvel movies.

I disagree on two fronts. First, given I responded to "> because authors prefer to be paid for their labor", I think the economics are the key, rather than the artistic merits. The authors and artists suffer economically purely on the basis of the AI doing their work for less, not on the basis of actual artistic merit.

Second, the subjectivity of artistic merit means different people like different things. From a market perspective, this is why horror films get made even though they disgust people like me, it's why kids films get made even though adults outnumber kids, and it's why RomComs exist despite the stereotype of men cringing at them.

You are however correct that it assumes no marketing exists to manipulate the best seller lists. But the marketing is also being outsourced to AI, and I suspect there were more writers writing copy than writing novels and screenplays. Now? Now I'm not so sure, though I'd guess it's still true.

I can sympathise with you about other people thinking you're just virtue signalling with your book collection. What I meant more along the lines of how War and Peace has a reputation for being a book that people like to claim to have read but actually have not, or how many loud atheists state that only atheists have read the bible and that's how they ended up being atheists, though in either case I don't know how accurate the reputations are.

If the data is available online for the pile, surely it's also publicly available to ordinary people in a way that means authors aren't getting any money.
What sort of defense is this? "Your honor, after someone broke in, they left the door open. Since the door was unlocked anyone could have committed the crime I'm accused of."
This is pretty reductive - "FTFY" is rarely the witty response you think it is.