macintosh.world | Log In | Register

Today | News | Books | Recipes
Notes | QuickTake | Wiki | Browse
Maps | Reference | Reddit | YouTube
Chat | Games | About

Back to HN

Can I opt out of my input or output data being used for training?

by teekert | 496 points | 243 comments | 2026-09-02 07:30:39 Central

Open Source Link | Read Source Here

Open on Hacker News

Comments

teekert
Context: After careful research our organization preferred
a European partner with good central privacy controls. We
landed on Mistral, after being disappointed that the Pro
tier was opt-in to training on prompts by default we
switched up to the Team tier which provides an
organization dashboard with some relevant settings. As we
did that Mistral changed these options and the Team tier
was now also opt-in by default and at the same time seemed
to have lost the ability to centrally disable training on
prompts for your entire organization. This even caused
some of our (testing) prompts to be used for training
(which Mistral removed after we expressed our
disappointment).For some time these pages conflicted with
what our users reported (they said that in contrast to
what I stated to our management they found they were opted
into training on prompts by default as per their own
privacy page). Mistral just now corrected their docs. I'm
not sure how long the conflicting situation has lasted,
but at least for several days.For contrast: Claude
disables training on prompts for organizations starting
from the 18 euro tier [0]. As a European I'm
disappointed.[0]
https://claude.com/pricing#team-&-enterprise

  > throwaway89201
> and at the same time seemed to have lost the ability
to centrally disable training on prompts for your
entire organizationThere is a toggle on
https://admin.mistral.ai that allows you to disable
training for both Vibe and Console/API for your entire
organisation. And I'm not on the enterprise plan. I've
disabled training the first time I created an account,
and it has remained that way.You story is also very
confusing due to the wording around "opt-in by
default" and "disappointed [about] opt-in to training"
(most people would be disappointed about an opt-out)
and probably conveys the wrong message to most people.

    > > teekert
I don't have that toggle (but could indeed have
sworn I saw it earlier).Sorry, I have always
thought that "opting in" is, "opting for the
presented option" and opting out is "opting out of
it", so opting out [of sharing prompts for
training] is choosing to not share, but apparently
I was wrong my whole life. I'm not a native
speaker, and I think most people here (in my
country) would interpret this the way I do? Weird
but TIL.

      > > > bee_rider
FWIW I don't find what you wrote confusing,
but I do think it is following a sort of...
bad convention that some companies have been
pushing.Opt-in and opt-out describe the nature
of the choice that you make. "Opt" means to
choose (apparently it is a French word we
stole). Opt-in means you have to proactively
choose to be in. Opt-out means you have to
proactively choose to be out. "Opt-in by
default" is an overly verbose way of saying
"opt-out."In either case it describes the
choice that you need to proactively make to
override the default behavior.Edit: I should
also say that it is a "known point of
contention" where pro-privacy people have been
pushing back on this phrasing. So, you have
probably accidentally stumbled into an ongoing
discussion, which is why some of the comments
might be unexpectedly prickly.

        > > > > rcxdude
>Edit: I should also say that it is a
"known point of contention" where
pro-privacy people have been pushing back
on this phrasing. So, you have probably
accidentally stumbled into an ongoing
discussion, which is why some of the
comments might be unexpectedly
prickly.Yes, because people will say that
something should be opt-in, meaning off by
default, and then someone makes it on by
default and says 'oh, it's opt-in, you're
just opted in by default!'

          > > > > > bee_rider
I 100% agree that "opt-in by default"
is bad and that "opted" in that sense
is a bullshit incoherent phrase.
Unfortunately this "opting as a thing
that happens to you instead of a
choice" idea has evidently been
promoted well, so we have to expect
that people will unknowingly use it.

            > > > > > > therealpygon
"Opt-in" means to agree to
something. If something is "opt-in
by default", it means that it is
opt-out and someone is using
weasel wording and contract
inducements to force you into
something you didn't actually
agree to. Any "disagreement" is
just people trying to advocate for
what they wished it meant.

            > > > > > > mahirsaid
And suddenly we have laws
directing AI companies to report
bad behavior or potential rival
candidates.

            > > > > > > half_fish
And some shadowy departments do
this without needing any laws at
all.

      > > > Miraltar
If you opt in it means that by default you're
not in and you chose that option. So if you're
now included in training by default, then it's
an opt-out feature as in you can opt out of
it.

        > > > > teekert
Idk, feels like my brain does not have the
model for this, this does not exist in my
language afaik haha.

          > > > > > bluefirebrand
Opt-in = "we assume you are out unless
you tell us you want to be in"Opt-out
= "we assume you are in unless you
tell us you want to be out"

          > > > > > sn0n
Just remember, good company's leave
them off by default. Bad companies
will turn them on for you.

          > > > > > blub
Replace opt with its equivalent
"choose" and it becomes
choose-(tobe-)in and
choose-(tobe-)out.

      > > > throwaway89201
> I don't have that toggle (but could indeed
have sworn I saw it earlier).You don't see
"Allow the use of your interactions with Vibe
to train Mistral's AI models" at
https://admin.mistral.ai/vibe/privacy ?And
"Allow the use of your API calls to train
Mistral's AI models" at
https://admin.mistral.ai/plateforme/privacy ?

        > > > > teekert
Mistral support just confirmed that the
org-wide toggle was removed from the Team
plan earlier this year and made exclusive
to the enterprise tier. I bet that if you
signed up earlier, they didn't remove it
from your account, but it's not there for
new customers. Which explains some of the
confusion here.Also explains that the
community support in Discord insisted that
the toggle should be there, moreover they
told me the individual toggles on the
user's pages wouldn't do anything because
it defaulted to "off" for any
organization, as per their docs until 2
days ago. But that changed.

        > > > > teekert
The toggles I seeat
https://admin.mistral.ai/vibe/privacy:Allo
w public sharing of chats contentAllow
user feedback on model responsesChat
Retention PolicyAt
https://admin.mistral.ai/plateforme/privac
y I seeAllow the use of your API calls to
train Mistral's AI models.Enable Labs
modelsSo indeed, no "Allow the use of your
interactions with Vibe to train Mistral's
AI models". I only see that setting here:
https://chat.mistral.ai/chat?profile_dialo
g=privacy
And I have to ask every user in my org to
go and turn it off.

        > > > > teekert
Ok it (org wide toggle) just (24h after
posting this) returned and the docs were
changed again! They now include the
text:"Vibe (Teams): Administrators can
disable data training usage for the entire
organization."That's nice! Except that it
was on for my entire org so I'll be
checking if they didn't store anything
first thing tomorrow. I wonder what
changed their minds.

      > > > duskdozer
No you're not wrong, but while I can't speak
for this specific instance, in many other
cases, this language is intentionally
confusing in order to dark-pattern users into
agreeing to the thing they don't want to.
Similar is using ambiguous language next to a
switch/checkbox that makes it difficult to
tell exactly what each state means.

      > > > abhayaya
Imho, legally, one should always be 'opted
out' of 'forced mingling of your ideas'
without deliberately opting in. Same as for
advertising, click-through tracking, and
sociopsycho/marketing metrics. Like the old
firewall rule (easier to handle opt-in/out
that way than firewalls; like, how many people
can run a firewall that way now?). Users
should have it right in front of their face
when they install something or register, too,
not hidden among options. It might take the
user five or ten more minutes to get to use
the app but worth it.

      > > > quadrifoliate
I think there is a simple, unambiguous way to
describe this, which is to use the passive
voice or an explicit subject with the past
tense to clarify that you didn't do the opting
in.Saying "we were disappointed to be opted in
to training by default" clarifies the point
you're trying to make, which is that the
toggle (whatever it may be called) was set to
training by someone else, not you.Actually in
your case you'd say "...after being
disappointed that the Pro tier opted us in to
training on prompts by default", which gives
an explicit subject ("the [Mistral] Pro
tier"). No one will confuse that with the
opposite "the Pro tier opted us out of
training on prompts by default".

        > > > > frereubu
I think the trouble with "opted in to
training by default" is that there was no
"opting", i.e. no choosing. I can't come
up with a better suggestion off the top of
my head though!

      > > > tedggh
You said it correctly, there's no confusion.
        > > > > tedggh
Downvoted for knowing how to read
    > > Barbing
My native-US-English speaker read:"disappointed
that the Pro tier was opted-in to training on
prompts by default" [and required manually opting
out]"the Team tier was now also opted-in by
default" [and required manually opting out]In
context of each sentence and the larger comment,
read smoothly here.Also - have seen more than one
lively discussion on these phrases, since defaults
can stick 95% of the time and Big Tech has done
their best to be abusive about what they
automatically enable for users by default for some
time.

      > > > rcxdude
It's understandable but it is technically
inconsistent and dilutes the meaning of
opt-in. Opt implies an active choice, so you
never are opting for the default.

        > > > > wtallis
It's also simply bad writing regardless of
what meaning the reader takes from it.
"Enabled by default" or "Allowed by
default" are much clearer and more natural
phrasings without any connotations of the
user having taken an action to express a
choice.

    > > qwertox
I also have these toggles, while being on a free
plan. I must have set them to disabled at setup
because that's how they are now.

  > lukan
"For contrast: Claude disables training on prompts for
organizations starting from the 18 euro tier "In
theory also for individuals?At least I have that
toggle to deactivate that. But how would I ever know
if they actually respect that?

    > > kccqzy
If you don't trust the company to keep their
promise, don't use their products.

      > > > RussianCow
I don't trust any company with their word on
anything. Luckily, privacy policies are
legally binding.

        > > > > lostlogin
> Luckily, privacy policies are legally
binding.Companies violate them all the
time and massive leaks happen a lot.The
punishments are trivial.

          > > > > > rkangel
GPDR fines in the EU are NOT trivial.
        > > > > tjwebbnorfolk
> Luckily, privacy policies are legally
binding.Laws are violated all the time.
The graveyard is full of people who had
the right-of-way at a crosswalk ...

        > > > > nullsanity
Why does that matter? Legally binding just
means "slightly more expensive when we get
caught"

      > > > LeBit
How can you trust any company?The only company
you could think of trusting is one where an
external , independent auditor is doing its
work.

      > > > locknitpicker
> If you don't trust the company to keep their
promise, don't use their products.Your comment
is perplexing. No company on earth meets your
requirement. What are you expected to do? Move
to a hut in the woods?

        > > > > dylan604
> Move to a hut in the woods?I'm so burned
out on the bullshit companies that succeed
in tech forcing their will upon us serfs
that I'm actively looking into that very
thing at this time. Tech used to be fun.
Now it's just depressing. I'm ready to go
back and take the blue pill.

          > > > > > lukan
Been there. Nice and peaceful, but can
get lonely, though.

      > > > roosterIllusi0n
That is ridiculous. Vote to ensure products
have to be legally private and make it a crime
to share or reuse your data for anything. This
should be the norm that people vote for. The
idea that we have to give up privacy so some
nerdy pervert can be a billionaire off the
backs of people who do the work is
absurd.Criminalize failure to keep private
data private. Arrest CEOs and executives. Put
them in jail when it happens.

      > > > lovich
What company do you trust to keep their
promise?

    > > teekert
This post is only about the defaults, and
expectations therefore for Team and Enterprise
plans and their organizational controls.

    > > danelski
My 20 EUR Claude (individual) had the training on
by default.

      > > > teekert
Yes, individual indeed (used to different but
changed about a year ago, see [0] among
others].For the Team and Enterprise plans de
default is "not sharing prompts for training"
(which I mistakenly referred to as opted out
of sharing prompts for training.) [1][0]
https://news.ycombinator.com/item?id=45076274[
1]
https://claude.com/pricing#team-&-enterprise

    > > warkdarrior
How do you know an LLM provider does not kill
puppies every time you send a prompt longer than
14 words?

      > > > lukan
Because they have no incentive as a company to
do that?
But do have a strong incentive to learn from
user input.
(Most cutting edge features are developed
private - very valuable to get that into your
tool)And getting the data is not hard, they
already have it.
Risky is indeed a bit making use of that data,
as that requires at least some humans (as
potential whistleblowers). But you don't even
have to tell them, where the data came
from.Whether they do it? No idea, I assume
not, but I see a risk.

        > > > > roosterIllusi0n
We can easily make it a crime to collect
the data and another crime to share it or
store it in a way that allows someone else
to copy it, authorized or not.People
always want to overcomplicate everything
to benefit 5 super rich people that got
rich taking wages from workers and ripping
off retail investors. Please stop the
pandering. Vote for yourself and your
neighbors, don't vote to eliminate privacy
so a rich pervert gets to exploit you.

          > > > > > lukan
And what would my vote change about
the situation?Local hosting will be a
solution, once the hardware becomes
affordable.

            > > > > > > roosterIllusi0n
Vote for people who want to
restore the 1950s federal tax
rates that we already had in the
1950s. I find so sad that people
ignore history and act like
solutions don't already exist.The
90% tax bracket in the 1950s
capped CEO and investor yearly
income to the 2026 equivalent of
5mil a year. Use that imaginative
brain that likes to speculate
instead of learning history to
image how our entire society
changes if no individual can earn
more than 5mil a year in all
income sources. We also banned
stock buybacks and had more
regulations put in place after the
1929 collapse and great
depression. Those things were
removed slowly over the 60s and
70s until a lot was gutted all at
once in the 80s. The final blows
came in the 00s.All the problems
that existed leading up to and
during the great depression are
back because we reverted all the
laws and policies that were put in
place to prevent it from happening
again.None of this is complicated.
Stop worrying about culture
nonsense or right wing nonsense.
Start voting to tax rich like we
did in the 1950s.Why did we have
so many local and regional stores
back then, but only a handful of
conglomerates today? There is no
incentive to merge companies by
CEOs when it turns two 5mil a year
CEO jobs into one 5mil a year CEO
job. In that environment, merging
is due to actual company need, not
CEO enrichment.CEOs will also stop
cutting wages, under staffing, and
outsourcing. It all stops because
they no longer get paid more doing
it. This is actual US history, the
country boomed economically
because of those tax rates and
financial regs. This is not
something anyone gets to deny, it
already happened.Voting against a
proven solution is madness.

    > > tedggh
Always assume and act as they won't honor it.
  > summarity
"Opt in by default" would mean that it is not enabled
by default. Do you mean opt out?

    > > kevincox
Yes, I found this comment very hard to understand
until I realised they were talking about it being
opt-out with no setting to change it. (At first I
thought it was opt-in by default, and the "by
default" implies that there is a setting to change
the opt-in/opt-out setting)

    > > teekert
You are "opting in to sharing your prompts for
training", by default in this case. My slider
says: "Allow the use of your interactions with
Vibe to train Mistral's AI models.", it is on by
default for everyone on the Team plan, the admin
can't centrally turn it off anymore, and any user
can toggle it when they want to. This all changed
last week.I know I'm naive but I expect that when
I pay, this stuff is simply off, so I was already
surprised by the Pro plan. But I did look out for
it there, because Anthropic made this switch some
time ago.

      > > > KPGv2
> You are "opting in to sharing your prompts
for training", by default in this case.The
English term for that is "opt out" not "opt
in." To opt is to choose. If something is on
by default, you have not opted in. You were
forced in, and turning it off means you must
opt out. (I.e., choose to be out.)normally I
wouldn't care about a mistake like this,
except that opt in/out are very important
concepts in software development and hacker
culture. And it reversed the meaning of the
original comment in a highly confusing,
relevant way.

        > > > > jrave
as another non-native speaker, i think
that while the person you're answering to
didn't use the default way of expressing
this in the english-speaking world, i did
understand what they meant, i think they
do know what opting means and i think the
reasoning is this:
when they say data collection is "opt in"
by default they mean that by choosing to
use a product, you are opting into your
data being collected (at the same time).
on the other hand, native speakers saying
something is "opt out" describes the
options or toggles one has available in
the default case - when something is
toggled "true", you (only) have the choice
to toggle it of.
so english speakers talk about the
controls one has to CHANGE the status quo.

          > > > > > kzrdude
As another non-native speaker, I
misunderstood what the person meant.

        > > > > brendoelfrendo
They said "opt in to training on prompts
by default," which is coherent English and
perfectly fine. The phrases "opt in" and
"opt out" will always be contextual based
on what is being opted, so it really falls
to the reader to pay attention to that
context. Please don't give English
language advice as though you are an
authority; there is not a rule of the
English language that would make their
usage unacceptable.

          > > > > > F3nd0
Since the whole idea of 'opt in' is
choosing to do something, the only
reasonable interpretation of 'opt in
by default' is that the default state
requires you to make the choice
yourself. However I try to bend my
brain, sing the term 'opt in' for the
exact opposite makes no sense
whatsoever.Note that I'm not using
English grammar as my argument; human
language (and English especially) has
plenty of exceptions and expressions
which make little sense, and if people
were to adopt this phrasing en masse,
it would naturally become a part of
the language, whether it makes sense
or not. My argument is that adopting
this phrasing is a terrible,
user-hostile decision and in my honest
opinion we really, really, really
ought to avoid it.

            > > > > > > PotatoPrime
I read it as though he was trying
to say '[you are] opt[ed] in by
default'.Now granted, he did not
say that, and I only reached that
conclusion by the rest of this
conversation, but I think to have
meant that and said '... was now
also opt-in by default...' isn't
unreasonable.Just sharing since
you mentioned you couldn't see how
using that term for the opposite
would make sense.

            > > > > > > F3nd0
Thank you. However, assuming you
mean 'an affirmative choice is
made for you' when you say 'you
are opted in', that use really
takes the 'opt' out of the
'opt-in'. You can opt in, but you
can't have someone else 'opt you
in' for you; that seems to go
against the very core meaning of
'opt-in' (i.e. having to
explicitly choose something
yourself).In other words, you can
explicitly choose something by
yourself, but if other people (in
this case companies, service
providers) choose something for
you, it's no longer your explicit
decision made by yourself. Hence
why I believe we can't consider
that 'opting in'.

            > > > > > > Timon3
Even read this way, it doesn't
make logical sense. It's a great
example of newspeak - corporations
know people use "opt-in" as a
purchasing signal, so they're
trying to redefine the word by
abusing it enough times."Opting
someone else in" is essentially
the same category of error as
"consenting for someone else".
People can only consent for others
in specific circumstances and even
then only for specific people
(e.g. legal guardians), trying to
argue "but I consented for them!"
in front of a judge usually ain't
gonna work.

            > > > > > > KPGv2
Here's the dictionary:
https://en.wiktionary.org/wiki/opt
-in> The property of having to
choose *explicitly* to join or
permit something; a decision
having the *default option being
exclusion or avoidance*; used
particularly with regard to
mailing lists and
advertising.You're just using the
term wrong. That's okay, I've used
words wrong before, and it just
means you've been presented an
opportunity for learning a new
fact.But there's no argument to be
had except with the dictionary.

            > > > > > > F3nd0
My use seems to be perfectly in
line with the given dictionary
definition, though? Did you
perhaps mean to address a
different comment?

          > > > > > mikebenfield
I find the usage bizarre and
confusing, as others also clearly do.
There's nothing wrong with pointing
this out.

          > > > > > KPGv2
> They said "opt in to training on
prompts by default," which is coherent
English and perfectly fine"Opt in" and
"by default" are contradictory
phrases. They mean literally the
opposite of each other. The only
reason it's coherent is that most
native speakers will elect to believe
"by default" was used correctly
because it's the easier to use term.>
there is not a rule of the English
language that would make their usage
unacceptableI think you're using
emotionally inflammatory language,
possibly unintentionally. no one is
refusing to accept what they
wrote.That being said, the way they
used "opt in" is categorically wrong.
It's the opposite of the dictionary
definition.https://en.wiktionary.org/w
iki/opt-in> The property of having to
choose explicitly to join or permit
something; a decision having the
default option being exclusion or
avoidance; used particularly with
regard to mailing lists and
advertising.

          > > > > > TulliusCicero
> They said "opt in to training on
prompts by default," which is coherent
English and perfectly fine.No, it's
really not. It should read "opted in
to training on prompts by
default"."Opt in" means "you have to
affirmatively turn this thing on".
"Opted in" means "the thing is
currently enabled".Which is why it's
probably better to write it as
"enabled by default". Much harder to
misinterpret.

      > > > nozzlegear
> You are "opting in to sharing your prompts
for training", by default in this case.I
understand what you're saying here, but maybe
"turned on by default" is less confusing for
everyone.

  > Exoristos
> opt-in by defaultThat's called opt-out.
    > > globular-toast
Yeah, it was confusing to read the parent before I
realised they got the terms the wrong way
around.To those wondering, "opt" means to choose.
"Opt in by default" makes no sense because you
didn't choose; this is just "in by default". If
they give you an option then it's called opt out.
Opt in would be "out by default" with the option
to go in.

  > rezonant
Buried by the distraction of the semantics of "opt in"
vs "opt out", the more interesting conflict isn't
addressed in the sibling threads.> and at the same
time seemed to have lost the ability to centrally
disable training on prompts for your entire
organizationvs> Vibe (Teams): Administrators can
disable data training usage for the entire
organization.Is the document out of date? Or did
Mistral reinstate this ability after the fact? What's
the story here? Others seem to be stating they have
and have had the ability to opt out of training
centrally for a long time.EDIT: Oh, it is addressed
just a bit hard to find with all the opt in/out
explanations:
https://news.ycombinator.com/item?id=49549102

    > > teekert
You are right! They must have just now switched
this back, I also have the toggle now! It was
turned on sadly but it's something.

  > teekert
Just now the sentence:"Vibe (Teams): Administrators
can disable data training usage for the entire
organization."Was added to tfa. And I now see an org
wide toggle where there was none before! Sadly it was
on so I hope no users submitted stuff in the mean
time, but it's something!

  > jacquesm
I don't trust any of these companies with my data, and
I assume that whatever data they've got is going to be
used, one way or another, no matter what they tell
you. It would be nice if you could stick a sentinel in
your data that if it ever shows up in the models you
know they've broken the rules for sure.

    > > bluefirebrand
Yeah. Unfortunately they know you can't catch them
on this so they feel absolutely free to do
whatever they pleaseIf such a data sentinel did
exist then we might see them change their behavior

  > Aldipower
Today I registered a free account with Grok, because I
simply was curious. Man, training is "off" by default
even with the free tier. I was positively surprised.
As a European I'm disappointed too.

    > > blazarquasar
Grok has possibly the worst ToS of any of the AI
providers.
They are probably different in the EU, but:> In
choosing to submit, create, generate, record,
post, or display Inputs on or through the Service,
you grant an irrevocable, perpetual, transferable,
sublicensable, royalty-free, and worldwide right
to SpaceXAI to use, copy, store, modify, process,
adapt, transmit, distribute, reproduce, publish,
upload, download, display in public forums, list
information regarding, make derivative works of,
and distribute such Content, including anything
referenced therein, in any and all media or
distribution methods now known or later developed,
for any purpose, and to aggregate your User
Content and derivative works thereof for any
purpose, including but not limited to: (i)
maintain and provide the Service; (ii) improve our
products and the Service and for our other
business purposes, such as data analysis, customer
and market research, developing new products or
features, or identifying or displaying usage or
User Content trends; and (iii) perform such other
actions to enforce these Terms, comply with our
Privacy Policy, comply with applicable law or
governmental, court, and law enforcement requests
or requirements or keep our Service safe.> To the
extent the User Content includes a person's image,
likeness, voice, or other similar attributes, you
grant SpaceXAI the same rights to use those
attributes as part of the User Content as
described above. You represent and warrant that
you have obtained all rights, licenses, notices,
permissions, and consents necessary for SpaceXAI
to use that User
Content.https://x.ai/legal/terms-of-service

    > > _puk
Lol, I'm fully expecting in 6 months time:"A bug
in our portal had the setting for training
inverted. This means that when you expected us not
to be training on your data, we actually were.
We know this adversely impacts the trust our users
invested in us, so as of today we are crediting
all affected accounts with $200 to use on our
latest models".

      > > > Aldipower
One can expect a lot what happens in 6 months.
Maybe the earth will turn clockwise then!
Think about it. :-D

        > > > > artwr
Lol. I would have thought it depended more
on your relative position to the equator
than on the season.

      > > > leonidasrup
Until the risk of large financial penalties or
long jail times is high enough, the CEOs of
there companies operate using the "It's Better
to Ask for Forgiveness Than Permission" model.

  > semiquaver
Sorry to nitpick but the scheme you're referring to is
called "opt-out."

  > kragen
> being disappointed that the Pro tier was opt-in to
training on prompts by default"Opt-in" means that the
default is non-participation, for example, not
training on your prompts. Is it possible that you
intended to say "opt-out", which means that the
default is participation? That's what the context
seems to suggest.See, for example,
https://termly.io/resources/articles/opt-in-vs-opt-out
/:> Data privacy laws like the GDPR and CCPA give
individuals the right to opt in or out of different
data processing activities.> · Opt in consent means
the user takes an action to show they agree to
something,> · Opt out consent is when they take an
action to say no.Or
https://bigid.com/blog/opt-in-vs-opt-out-consent/:>
• Opt-in consent requires users to actively agree
before data collection or processing.> • Opt-out
consent allows data collection by default unless the
user declines.This is an important distinction,
because confusing the two (as you seem to be doing)
can lead you both into unethical fraud and legal
liability.

20k
You have to be rather naive if you don't think these
companies don't simply train on your prompts with or
without your consent. They literally scrape everything -
legal or not - and claim its fair use to train on,
including straight piracyThe idea that they'll steal from
everyone except you is just wishful thinking

  > teeray
This is why these companies paying subscriptions
shoveling everything into Claude thinking "oh, they're
not training on our stuff" is hilarious to me. Of
course they are. They probably are extra super-duper
sure not to have Claude admit to that in any way, but
there's no way they're giving up training on the sum
total of both open and closed source code out there.

    > > vorticalbox
Good example of this is figma. Claude was updated
to work with it and now we have Claude design.

  > protocolture
>You have to be rather naive if you don't think these
companies don't simply train on your prompts with or
without your consent. They literally scrape everything
- legal or not - and claim its fair use to train on,
including straight piracyHave been having a think
about this statement. I agree in intent, some of them
are probably breaking the agreement for training data.
I dont think Microsoft is doing it, Enterprise Data
Protection is the plank holding up their entire
Copilot line. One whiff and everyone's gone. Copilot
isnt actually good at anything except giving some
illusion of protection, and preventing users from
following a desire path to other LLMs without
enterprise data protection. Its the core value
proposition. But we only need to wait and see what the
next 20 data breaches tell us to find out for sure.In
detail however, I dont know if they could just claim
it as fair use, after exclaiming that they
specifically wont do that. I dont think "Fair Use"
would be the issue so much as contract law. I know a
EULA wouldnt hold up the other way (By reading this
you agree not to steal my data and train an LLM with
it) for data thats made freely available on the
internet. But if you are purchasing the "No Training"
contract they would be in breach if they trained with
it. Possibly fair use would let them keep the data
after paying whatever they owe in terms of contract
breach.

  > rcxdude
This only makes sense if your model is 'the labs are
just breaking the rules all the time' as opposed to
'the labs have a legal theory of defense for the one
aspect of the business which is legally questionable'.
The question of whether doing something their terms of
service explicitly say they will not do opens them up
to civil liability is a much more certain one than
anything they are doing around copyright. If they are
doing this and anyone can prove it, they are going to
be sued very hard, and it will be a pretty
straightforward case.(To put it another way: if the
copyright arguments against training on data scraped
from the internet fail, the big AI companies don't
really have a business, so they are going to proceed
on the basis that they do until someone forces the
issue otherwise. They don't need to train on data from
their customers, and that is an argument that is
almost certain to fail in court. If you're carrying
100kg of cocaine in your car, speeding is a really
dumb idea, but the analogous action here would be
firing a machine gun into the air from the driver's
seat)

    > > AlexandrB
It doesn't matter. Even if they don't train on
your inputs now they will "boil the frog" and do
so in the future. What the user wants is
irrelevant in the tech industry.> If they are
doing this and anyone can prove it, they are going
to be sued very hard, and it will be a pretty
straightforward case.This is a joke, right? The
usual settlement in these cases amounts to a few
days worth of revenue.

    > > knollimar
Well maybe they have another suspicipus theory for
defense of your data. Like taking it, laundering
it through a summarizer, and training on that.
Surely that's less questionable than taking
copyrighted stuff and saying "don't output over 15
words, never verbatim" pretty please.

      > > > rcxdude
Copyright has a lot of leeway for arguments
about fair use. 'We won't use your data for
training' in a contract has far fewer.

  > gitgud
If you have an enterprise contract with them, the
legal protections for the consumer is much higher...
is what I've been told anyway

    > > 20k
It isn't, in both cases you're protected by
exactly the same legal system

      > > > zamadatix
It's also the same planet but finding a
commonality somewhere in the chain is not the
same thing as finding a lack of difference in
the matter.

      > > > gitgud
Same legal system, but an OpenAI enterprise
contract is not the same as a ChatGPT pro
subscription... The terms of service agreement
would be quite different

    > > olejorgenb
Not sure what you are saying exactly, but> the
legal protections for the consumer is much
higherYou know you give away the right to file
class action law suits against Anthropic when you
accept their Terms Of Use, right? (at least the
Americans ones)

      > > > stingraycharles
Is that the case with enterprise contracts as
well? Can't imagine any decent procurement /
legal team accepting this.

    > > autoexec
If you have an enterprise contract with them
you're still unlikely to ever know that you've
been lied to unless some whistleblower at the AI
company comes forward and even if you do somehow
find out, it's too late. Once they have your data
and have trained on it there's no taking it back.
At most they'll pay out some tiny settlement
that's a fraction of how much money they make in a
week and they'll continue to profit from your data
forever.

  > j4k0bfr
I think innocent-until-proven-guilty is the correct
approach in general but... It's also immature to
assume your data will stay private in the long
term.These AI companies are immensely valuable
targets. When they get breached, their datasets will
inevitably penetrate public datasets. And why would
any AI company refuse to train on 'public' data?

    > > PunchyHamster
innocent until proven guilty is definitely wrong
approach vs tech giants who time and time again
prove they will do anything unethical if only they
can theoretize how to get away with it or the fine
is low enough

  > addag
And even if they are not doing it right now, they
probably keep the history available for future
training, "just in case".

  > WarmWash
It would be catastrohic for any of the big labs if it
came out that they were training on what was sold as
private.I get this cynical conspiratorial energy, it
fits the internet well, but I can assure you most
people with even mild business sense would be
intensely opposed to this idea. Well, except maybe
Zuckerburg, but they don't really do enterprise
anyway.

    > > thephyber
No it wouldn't.It wasn't "catastrophic" for the
largest of the 3 US credit reporting agencies when
their entire dataset was breached. The company is
100% IP and the only value they have was
completely copied. Their largest value is to
verify identities by the things Americans know
(KDB) and after that "single factor of identity"
was 100% compromised, the company only got bigger
and more contracts.When there are only 4
competitors in the large scale foundation model
business and they all throw caution to the wind
because they are racing to own the "$30 trillion
TAM" they are all going to make critical security,
RBAC, and segregation mistakes.Both ChatGPT and
Claude threads marked for sharing have been
indexed in Google at large scale. This is
incredibly easy to tell Google crawlers via
robots.txt not to crawl those URLs, but nobody at
either of these uber unicorns could be bothered to
add that one pattern to the one file.And all of
the skepticism here is about verifiability. The
foundation model companies are liable for
potentially more the companies are worth if found
to be violating copyrights of content used for
training. They aren't going to make it easier for
lawsuits against them by detailing their data
ingestion into training pipeline.

      > > > avianlyric
The credit reporting agencies lost the data of
normal people, they didn't lose their
customers proprietary internal data. The
credit agencies didn't loose or misplace their
customers data, so obviously their customers
don't really care that much, and the credit
agencies weren't sued into oblivion.But I can
guarantee you that if a companies internal
data got leaked or misused, then every single
enterprise customer of that lab would turn
around and start suing them. As an enterprise
customer you would be foolish not to, if only
if figure out via discovery just how badly you
got screwed.You want to see how nasty that can
get. Just go and look at what Apple is doing
to OpenAI at the moment. Do you really think
Apple wouldn't find a way to sue a lab into
oblivion if they discovered a lab had secretly
started training on their data?

      > > > jppittma
that was accidental? this would be straight up
fraud.

        > > > > chillfox
You can structure it so that it becomes
accidental. 1. Ensure security barriers
are weak or honor based.
2. Put individual researchers under a lot
of pressure.
3. If you get caught, blame the weak
barriers, or the individual researcher.

Basically setup the incentive structure to
incentivize researchers sticking their
mittens in the private cookie jar while
putting the cookie jar in a dark
unmonitored/unsecured room with a sign on
the door saying please don't enter.

          > > > > > thephyber
Yeah, the steps follow exactly what
happened at VW with DieselGate. The
diesel emissions lies were found out
because some enterprising person set
up an emissions testing system and
drove the car in real world scenarios
with it to verify the claimed
emissions.There's no reliable way to
verify a foundation model has been
trained on a particular piece of
proprietary data. If an API key is
ingested, hopefully the foundation
model is wrapped in enough moderation
that the raw API key oberserved during
training is not recited verbatim in
the output.

    > > r_lee
exactly, I don't understand how HN doesn't
understand thisby this logic every business
contract in tech is just a bunch of lies and means
nothing and the only way to do anything is to have
a server sitting next to you, otherwise it's
"someone else's computer"

      > > > 20k
I mean, they've already violated the law in
acquiring all their training data already, why
would they be uncomfortable violating a
contract to get more training data?

        > > > > vineyardmike
Because violating the law was the
prerequisite to starting their business,
without it they're worth $0 and have no
models. They've already survived Training
on customer input may help the models but
it isn't "bet the farm" helpful. Now, they
have a thriving business, so they
shouldn't risk their business for
incremental data that they can buy.At this
point, the reputation of the business
matters too. Fable explicitly didn't
support zero-retention usage, and it saw
significantly lower adoption vs other
flagship models, and their past releases.
Being caught abusing enterprise contracts
is really hard to dig out of.

          > > > > > r_lee
and they're under no contract with
those millions of websites and books,
but it's a whole other thing when it's
a paying customer, especially a large
enterprise with a signed contract with
DPAs etc.

    > > realusername
It's going to be the same as PRISM, people will be
outraged and business will go back as usual.(And
they totally won't do it again they swear, the
contract says so)

    > > ajb
It would be catastrophic if they violated
confidentiality blatantly, but that doesn't
exclude learning of any description. For example,
an ordinary human being can't fork a subagent for
a particular client and wipe it afterwards. Humans
can't stop themselves learning, so confidentiality
can't ban all learning.Instead, confidentiality
includes not literally copying material, not using
trade secrets or inventions, and not using
knowledge of business dealings for your own
purposes.So, while I fully expect that the big
labs don't train on private material to the extent
that they do those things in a blatant way, it
would not be surprising if they pushed the
boundaries. Humans push the boundaries all the
time.Up until now, machines did not have
judgement, so if you set up a machine in such a
way that you hadn't ensured it couldn't violate
contract, you were culpable. But now that they
have some kind of judgement, maybe it's enough to
avoid liability to tell it to obey the contract,
even if you give it incentives not to. After all,
that's how it works with human employees, isn't
it?Perhaps now we have machines that understand
language, someone somewhere is working on getting
them to understand "a nod and a wink" as well.

    > > tripledry
I think it's arguable that it wouldn't be
catastrophic.Doesn't have to be outright lying, it
can be something like "opt out" being off by
default, and them opting you in at the next update
without you noticing.This is just a gut feeling
and goes into the conspiratorial energy, but I'm
pretty sure I can find countless examples of
companies behaving like this in the past where it
wasn't catastrophic (google meta amazon adobe ...)

    > > 20k
I mean the models have literally trained on:1.
Child porn2. Stolen music3. Private github repos,
before that was 'stopped'4. Illegally pirated
booksThem training on company prompts against the
terms of service would be one of the least bad
things that these companies have trained AI models
onWhy do you think a company - willing to break
the law for child porn - won't break the law when
it comes to your personal data?

      > > > Henchman21
These folks need to feel repercussions so hard
their souls flee to the afterlife leaving only
their sad, dead husks behind.

  > solaire_oa
Pinky-promises, embarrassing. I'd point to tinfoil.sh.
I'm not a shill for tinfoil, I haven't even used it or
looked past its homepage really, but if we are able to
legitimately secure privacy, then training data
questions are moot (and the market for training data
would probably shift/expose).

    > > lukewarm707
agreed. the pricing is a nightmare. either $20 for
codex and up to 00 millions of tokens or $0.50
cache in per million in tee.

  > someothherguyy
https://beckreedriden.com/the-black-box-problem-in-ai-
trade-...

  > classified
I really don't get how there could be anyone to whom
this is not beyond obvious.And yet, software companies
are donating their means of production to these crooks
left and right, hastening the day where they really
will become obsolete. I'm sure others do equally
unwise things.If you're big enough, crime pays
exceedingly well.

  > deadbabe
I don't care about what they do, I only care about
what they say, that way, when compliance requires us
to use LLMs that don't train on inputs, I can point to
that policy and continue on using the LLM.If they are
secretly scraping input prompts, that no longer
becomes my problem. Someday, a massive lawsuit comes
down, and everyone gets to play the victim.You're
naive for thinking we don't know how this really
works.