edot Tom Dietterich (Editor in Chief at arXiv) posted on
LinkedIn the other day that they're having trouble keeping
up with the onslaught of AI-generated papers. Lots of
suggestions in the comments but no magic bullets.It's
ironic that the LLMs which benefit so much from reading
arXiv papers of yore are now being used to pollute it.
Personally if I see a single author post 2023, I assume
it's junk, especially if they're not from an actual
research institution. Not all solo independent researchers
are phonies but ... many phonies are solo independent
researchers.OpenReview is okay but even some of the
reviewers are apparently using LLMs or just hardly
reading.
|
> bobmarleybiceps A recent review of mine had a LLM-ism at the very end,
"would you like me to format this into a formal peer
review report?" So they very likely copy-pasted their
whole review :). I'm pretty down on academia atm ;-;
|
> > coolness Super unfortunate. On the other side, I reviewed
four papers for a top-tier AI conference and three
were clearly fully Claude generated, as in all
text, figures, results, everything. Actual good
reviewer time is wasted on such papers and your (i
hope) human written good paper receives AI
responses. It's a sad state of affairs for sure.
|
> > > stingraycharles I find it incredibly sad that this is what the
world is becoming, and I mostly blame people
for this, not the LLMs.It's the same type of
people that would have no issue letting an LLM
open a pull request on GitHub wasting valuable
time of other humans, and whatnot.I'm using
LLMs all the time myself, but it's so
incredibly important to use it to improve the
quality of your work, not degrade it. People
seem to be totally oblivious about this.On the
flip side, it does make it easier to recognize
people who are wasting my time.
|
> > Skyy93 I think this is more of a systematic issue. I
review now since 2-3 years, I do not get paid,
which is fine. However, it takes always a huge
amount of time without really having anything from
it, but I do it because it is important work.There
now so many researcher that need to publish which
explains the flooding, LLM only speed it up, so
reciprocal reviews take place. So now you are
forced to review and you are having less and less
time. So it's a natural choice for you if you
already took an LLM to write a paper to use it to
review.Perhaps one solution would be much harder
entry barriers, and enforcing some guidelines. For
example that a supervisor can not have more than 5
papers and PhD students only need one real paper
on a major conference/journal.
|
> > baq You'd better not look at the average CI pipeline
in software shops then
|
> > cobbzilla someone was a meat proxy
|
> > > coolness https://gruhn.me/blog/2026-08-03/ in case
someone missed the reference
|
> > > > pbronez Oh my god that was just written at the
beginning of August?? It feels like I read
it a year ago. We really are speed running
this tech cycle...
|
> auggierose If you see any post on arXiv you should assume it is
junk.
I find it ridiculous that people put any value on
something being posted on arXiv. That doesn't mean the
post is bad. It just means you need to find other
means of judging it, for example by actually reading
it.
|
> > gus_massa I half agree. It used to be good, they called them
"preprints" because they were already sent to a
journal and somewhat expected to be accepted.
Until people noticed that they were not forced to
publish the "preprint" later so it got flooded
with crap, hand crafted artisanal crap.Unless you
are working in the area of the paper AND know the
reputation of the authors AND take a deep look,
just give it the same credibility than to a random
PDF posted in WordPress. They have some weak
filtering because to post in the arXiv someone
must vouch for you or something similar, but it's
a very weak filter and people was already abusing
it.And then the AI slop truck hit...
|
perdy Published there in August, from a company rather than a
university. Without arXiv we'd have had a blog post and
nothing citable.
|
> flexagoon A blog post is just as "citable" as an arXiv preprint.
There's no fundamental difference between them.
|
> > emil-lp One is immutable with a DOI.
|
> frumiousirc > PublishedPosted. Having a preprint on arxiv is not
publishing.
|
aurareturn Are they doing to do something about authors using Arxiv
to publish propaganda/opinion pieces but presented as
research?ArXiv gives the appearance of scientific
credibility that a blog post wouldn't have so I'm seeing
the platform get abused.Example:
https://news.ycombinator.com/item?id=49580164
|
> dooglius Probably not because the whole point is they don't do
critical review> ArXiv gives the appearance of
scientific credibility that a blog post wouldn't
haveThat's an error on your side not theirs
|
networkOne Sorely, sorely needed.How is science to evolve if good
research requires $49 a pop to view?And what amazes me is
that these authors have undeniably stood on the shoulders
of giants in order to create their research.
|
> kadoban You think authors are the ones driving the $49 view
fees? They are not.
|
> > SonicTheSith They are, if their institution does not have
contracts for that specific domain from Elsevier
et al.At my institute we for example do not have
access to parts of Springer.Luckily, more and more
is moving towards open access.
|
upupupandaway Asking earnestly: is ArXiv a valuable resource? I used it
few years back when finishing my (late) Master's, but also
saw a lot of crap published there (primarily for
promotion/visa purposes) so it kind of took the shine out
of the service for me.
|
> txhwind A free-to-read PDF host site is valuable enough. Most
academic publishers have a pay wall.
|
> lemontheme My take: if not for arxiv and huggingface, the field
of ML would be nowhere near where it is today.More to
your question, I recommend something like
semanticscholar to find actual relevant papers. Try to
identify researchers that seem trustworthy, then
explore the citation network around them.For more hot
off the press stuff, follow what gets boosted on
social media.Still doesn't cover the truly niche stuff
but it's a start
|
> tristanj The preprint papers on arXiv are like 99% the same as
the published versions, except they're free instead of
locked behind a multi-thousand dollar/year paywall.
|
> > upbeat_general And if there are differences, that is often a good
thing! It can mean the author wanted to format
something in a particular way that the journal
didn't allow.
|
> > > senderista Often they are extended versions of the
journal articles and contain valuable material
that had to be cut for space limits.
|
> ufo Some researchers and journals publish peer-reviewed
work on arxiv; it serves as a stable archive that
won't be plagued by link rot or paywalls.
|
> edot It's good for getting free access to preprints which
are often close enough to the paywalled real papers in
journals. If I find a paper I want (or more
realistically, if ChatGPT finds a paper it wants for
me), but it's behind a paywall, odds are the authors
put a preprint on arXiv.
|
> embedding-shape Either you setup feeds for the specific
topics/subjects you care about, scan what you come
across once a week, or you use it to get access to
papers that are usually behind some paywall. I don't
think the intention nor the value comes from just
haphazardously reading through everything in some
section.It's not peer-reviewed and supposed to free
and accessible from both sides so the results kind of
makes sense.
|