Mathematicians want proof OpenAI didn't use their work

(theverge.com)

45 points | by kevcampb 1 hour ago

15 comments

  • sobiolite 26 minutes ago
    We shouldn't forget that the only way this unpublished research is supposedly getting into the training data in the first place is because the researchers making the accusations were conversing with ChatGPT in the process of doing their own research on these problems.

    But if they are saying this contaminated the model's training data with knowledge of their ideas, who's to say the model they were developing their research ideas with wasn't already contaminated through prior discussions with other researchers about the same topics?

    So if contamination is proved, or cannot be disproved, and if these researchers want OpenAI to relinquish its claim to have solved these problems independently, then it would seem they also have to give up their claim to have solved them independently?

    • loose-cannon 16 minutes ago
      1. Tristan's research direction was known to only a handful of other academics. It might help if you can be more specific regarding the origin of the contamination.

      2. If someone did come out and claim that their ideas were used without proper attribution in Tristan's work, then of course, that deserves consideration.

      3. What you're suggesting seems purely hypothetical. At present, there is nobody claiming that Tristan's work is "contaminated"

      4. Tristan was very willing in his initial statement to give credit to the people who developed the ideas.

    • glimshe 16 minutes ago
      This. I don't like people using others' work without giving proper credit. At the same time, people are jumping too quickly to defend these mathematicians without any solid evidence of their claims, while framing this as some sort of "Evil AI Company vs human mathematicians".

      In the best case scenario, the mathematicians were standing on the shoulders of the extensive training data from sources that aren't being credited and they may not even have had access to.

      I have nothing against these mathematicians because I don't know them. And given that, there's no reason to trust their word any more than OpenAI's.

      It's possible the work they were doing, even if related, was a dead end and immaterial to OpenAI's findings. Or maybe they are right and OpenAI stole their work. Who knows the truth right now?

      In other words, we need more evidence before making accusations.

    • api 23 minutes ago
      So it’s Napster for ideas?
  • eggy 31 minutes ago
    If you run a model locally to work on your area of expertise in mathematics, and you make a discovery, haven't you stood on the shoulders of others that created that model you are using? Playing devil's advocate here, are we only in for human collaboration, but not machine contributions in this case? Scholars get cited, but do they normally get paid for others citing them and using their work to advance their own? I get the whole sneakiness about how these companies like OpenAI built their models upon IP and data without a clear trail of attribution or compensation where it would normally be present. I am not a proponent for either side at the moment. I am trying to grapple with this whole new world of shared knowledge and how it is produced and shared and profited from especially when there is a $1m prize to distribute!
    • jaccola 26 minutes ago
      There’s always been a distinction between published work and unpublished.

      If my friend showed me some unpublished work (Say 90% of the hardest work toward a significant proof) if I then go finish the remaining 10% and publish the whole thing as mine it isn’t ‘building on the shoulders of giants’ it’s plagiarism/theft.

      If he’d published that work first and I took it and found another proof and credited him for his work via citation then that would be fine.

  • arutar 37 minutes ago
    Here is some (adjacent, but relevant) context, which has been posted elsewhere but I think is worthwhile to mention again here:

    https://mathstodon.xyz/@tao/117237320796901560

    Especially in recent years the mathematics community has worked very much in good faith, and a lot of effort is spent trying to give appropriate credit for ideas. Even when ideas are discovered in parallel if it turns out that previous work contained the same essential idea it is by and far regarded as best practice to give priority in this case. The point, in good academic practice, is to maintain the health of the practice at large.

    As Terence Tao explains in that post, the goal of mathematics is not only to solve big problems. And, even if one were very single-mindedly focused on solving big problems, it is still (in the long term) better to maintain the health of the community at large so that problems which are out of reach at the moment may be in reach again in the future. Good academic practice is one part of this culture.

  • asimpletune 24 minutes ago
    It's a little like we've gone back to the problem of the customer becoming the product.

    If I were a mathematician I would not my unpublished work to go into the hands of a competitor.

    If I were a lawyer I wouldn't want private details of my defense to be made available to the prosecution. Anonymous or otherwise.

    I wouldn't want the plot to an unreleased book to be suggested to another author.

  • throwaway713 40 minutes ago
    Am I missing something obvious? Isn’t this just a simple DB query to see the state history of the “Data Controls” → “Improve model for everyone” toggle in the settings? Just report whether that was ever on and over what time period.
    • heaney-555 36 minutes ago
      Multiple OpenAI staff have publicly said they cannot do that as accessing specific user settings without their consent (or legal requirement) violates their internal privacy policy.

      However, the mathematicians could easily declare whether they had the toggle on or off. Yet curiously, they will not say!

      • LelouBil 30 minutes ago
        The reply from OpenAI should have been "if you had the setting on yes, if not then no".

        It was very weird on it's own almost like the setting didn't matter. Or the researcher was being dishonest with what they shared to The Verge

    • monster_truck 24 minutes ago
      Yes you're missing several obvious things. Even saving the last changed date (nevermind every change date or what the change was) for every setting for every user would be earth crushingly wasteful. The value by itself isn't even worth including in backups.
  • aennassiri 32 minutes ago
    If OpenAI remained a full nonprofit looking to build an "OPEN" AI for the benefit of all humanity (not only the US or a few shareholders), I would have been happy to share my code, my work, and even label their data... This said, I don't blame them. It's a difficult mission to remain a nonprofit and, at the same time, have the required capital investment to build AGI.

    I'm not criticizing them, but I hope this race towards the first-best result or AGI doesn't blind them to making good decisions such as not using their users' data without consent.

  • faangguyindia 29 minutes ago
    I was trying to patent our maintenance tracking algorithm, which produces guaranteed weight loss or gain within 2–4 weeks by producing accurate calorie and macro targets for people to follow; in our test, it beats GLP-1s like Ozempic, Tirzepatide, and Retatrutide in results.

    But later we found that algorithm and math cannot be patented.

  • monster_truck 28 minutes ago
    I get the impression they have been LLM-psychosis'd into believing this
  • trescenzi 51 minutes ago
    If I understand correctly OpenAI cannot provide it. Because the models are essentially black boxes, especially this far after the fact, determining if this result built on training data based on conversations about the problem is impossible. So unless they can prove those conversations were never used for training then there’s no way to know.
    • yturijea 45 minutes ago
      It might be a lost cause regardless, because this is under the assumption that we can trust OpenAI to be honest about their own investigation, which is unlikely.
    • fractorial 41 minutes ago
      Only under gross negligence would it be unprovable: Did you use a model whose training set included user data? Did the transcripts of any of the agents include a tool call whose result including user data?
      • red75prime 15 minutes ago
        Or under strict privacy measures, maybe? When you are forbidden to connect a user and the user's data.
      • heaney-555 35 minutes ago
        >Did you use a model whose training set included user data?

        OpenAI, as with all AI companies, openly admits that it trains on user data unless the user opts out. But the mathematicians have not said whether or not they opted out.

        >Did the transcripts of any of the agents include a tool call whose result including user data?

        They have already explicitly denied this.

    • xxs 38 minutes ago
      It's the 'conversations' of some mathematicians with OpenAI. So the question is: did anyone feed the conversation(s) back as training data.
    • worldsavior 38 minutes ago
      What? Can they just look if they fetched certain documents/conversations and feeded them into the training loop?
      • heaney-555 32 minutes ago
        You seem to be confusing context-fetching with training.
    • heaney-555 45 minutes ago
      This is exactly the issue.

      What the mathematicians could do is reveal whether they had the data-sharing opt-out on or not. But curiously, as far as I've seen, none of them will answer that question!

      • socialcommenter 35 minutes ago
        Andreas opted out on June 29th[0]. As discussed elsewhere on HN[1]

        [0] https://mathstodon.xyz/@andreasthom/117240535270608201

        [1] https://news.ycombinator.com/item?id=49638353

        • heaney-555 33 minutes ago
          Then none of his work after June 29 will be included in the training data.

          What I was referring to is the fact that neither Levent Alpöge nor Tristan Buckmaster will answer this question.

      • aenis 39 minutes ago
        And its reasonable to assume that if they did, in fact, opt out, they'd make it clear.

        Most people do not understand that the main reason for the subscriptions is to give OpenAI and Anthropic the priceless, unique data that shows how the models are used, what people are building, how they are building, which solutions they consider OK, which they consider bad -- they purchase this data with cheap tokens. This is their only moat, really. If some really proprietary IP gets swept in the training data set its not really OpenAI's fault -- its the researchers'. Have something secretive? Dont fricking paste this into chatgpt. Duh!

        (I'd definitely not think OpenAI/Anthropic ignore the opt outs, or ZDRs. All it would take is one whistleblower to get them into terminal troubles. And why would they do it? They are not in the business of scooping unique IP -- they are in the business of understanding how AI is used across a variety of mundane, day to day work of individuals and companies. Useless math problem is good (or bad, as in this case) PR, but otherwise entirely worthless for the labs.

      • cmiles8 43 minutes ago
        Well that assumes OpenAI respects that setting. Given some of the company’s ethical challenges to date that’s not something folks are willing to just assume is happening.
        • brookst 39 minutes ago
          How does OpenAI’s respect (or lack thereof) for the setting change whether the people involved could say whether they had opted out of training? They comment you replied to noted that none of them had shared that info. How are they blocked from doing so?
        • heaney-555 38 minutes ago
          If you could prove that, it would be a gargantuan class-action lawsuit. And there would almost certainly be at least one whistleblower.
      • asimpletune 34 minutes ago
        One of them said they turned it off in June, back in the original mastodon thread.
        • heaney-555 33 minutes ago
          Then none of his work after June will be included in the training data.

          What I was referring to is the fact that neither Levent Alpöge nor Tristan Buckmaster will answer this question.

  • fooo1882992 25 minutes ago
    Looks like OpenAI's PR department woke up and put their main man heaney-555 to work on shaping public opinion.

    Dude makes up like 50% of the replies here.

  • kova12 35 minutes ago
    Do I understand it right that people now claim an ownership of the actual mathematical methods? What next, people patenting the letters and the numbers? And then the sounds? This starts going ridiculous.
    • loose-cannon 32 minutes ago
      dude, proper attribution is part of scholarship and academics. Not giving credit to the people whose work led up to your solution is malpractice.
      • inglor_cz 7 minutes ago
        A lot depends on the concreteness/vagueness of the idea.

        "Just use some polynomial" would be an idea so generic that it wouldn't merit attribution. "Just use a cyclotomic polynomial of degree 24" is a lot more concrete.

        (I have no idea where the exact limit lies and I have no idea how precise or vague were the ideas in those discussions.)

  • scotty79 38 minutes ago
    Math is excluded from copyright. So you can use any piece of math you ever heard from anyone and publish it, whatever the context, I think.
  • krater23 18 minutes ago
    When you want your unpublished work unpublished, don't store them on other peoples computers. Especially not on people their job it is to use data to create money.

    Would this data moved through a hack to ChatGPT, this would be another thing, but like this. No pity at all.

  • sylware 37 minutes ago
    Huh?

    Re-using advances done by others is necessary in maths.

    This is how hard sciences do progress.

    • xxs 35 minutes ago
      undisclosed ones in this case
      • sylware 31 minutes ago
        What's they were hacked and their current work in progress 'stolen'?

        On my open source projects, there is always a big "WIP" messy phase which I don't "really" publish... because it is messy.

        • xxs 15 minutes ago
          > they were hacked and their current work in progress 'stolen'

          in most jurisdictions it'd constitute a crime.

  • forlorn 47 minutes ago
    Why are people so possessive? Why wouldn't they just release their work for the profit of humanity? I'm genuinely curious.
    • fxwin 43 minutes ago
      Because in academia, your name and your body of work is a big deal that can open or close doors, and assigning credit is an important part of it. This isn't a new thing/exclusive to the AI era either.
    • cyclopeanutopia 42 minutes ago
      I have the same question, why OpenAI and Anthropic just don't release everything as open source?
    • kedikedi 44 minutes ago
      I don’t think it is people being possessive but rather people finding their value over the work they do (sentence ended up being weird but I hope the idea is clear).

      So I see someone dedicating a sizeable portion of their life on a discovery and not being credited is the problem here. Otherwise it would be published in a journal anyway.

      • krater23 7 minutes ago
        When you put a unpublished sizeable portion of your life in a ChatGPT window, then it's your own fault. I trust these AI companies as far as I can throw one of their data centers.

        When your work is unimportant enough for you to not setup a local AI, just shut up.

    • pluc 43 minutes ago
      Because it's not humanity that profits, it's OpenAI.
      • brookst 41 minutes ago
        Walk me through how the sofic result OpenAI publixhed profits OpenAI but had no benefit to humanity? Why were Thom and others working on it if it had no value?
        • pluc 38 minutes ago
          We've had the end result all along without OpenAI. All this technology does is, for a subscription fee, adds a middleman who you can talk to naturally. How does that which solely uses all the technology preceding it, is by itself a benefit to humanity? Get your head out of the stock market, AI is by design not creating anything new.
        • halsafar 28 minutes ago
          Very simply put it's an advertisement for OpenAI.

          Just like their "oops we hacked hugging face".

          Everything OpenAI does is to drive hype. They aren't honest or a good actor but they sure have good hype.

    • xxs 42 minutes ago
      a quote from the article: "Thom said it “would be ethically indefensible” if nonpublic research supplied by users helped to improve models that the company then used to race those very same users to publication, without consent, proper disclosure, or credit."
      • krater23 4 minutes ago
        Would one of this users tell their opponents what he is doing? And if yes, would anyone wonder when the opponent would use this information. Be dumb, win dumb prices.
    • diydsp 40 minutes ago
      1. Academic/research credit

      2. There was a $1,000,000 reward

    • timthelion 27 minutes ago
      OpenAI is the possessive actor here. They were supposed to develop the AI for the benefit of humanity. Now they do it closed and do it only for their private profit.
    • LelouBil 39 minutes ago
      There was an other thread on HN where apparently OpenAI promised one of the author half of the prize money IF they remove their co-author (that works for Anthropic) from their paper.

      This kind of bribing attempt doesn't seem to come from the "good side"

    • brookst 43 minutes ago
      Not defending it but in academia credit and attribution can be big deals. Some people are motivated by money, some by fame, some by satisfaction from doing the work, some by recognition, most by some mix.
    • bryanrasmussen 39 minutes ago
      As I lay dying here in this basement surrounded by my art, without food or basic necessities as I have given up all thought of material comfort in hopes of bettering the lives of others, I too wonder why others do not follow my example and live with no thought of their own comforts or position in society and strive to do things that may somehow benefit the species as a whole, especially the Billionaire part of that species.
    • suddenlybananas 43 minutes ago
      Mathematicians are way more open about what they do than OpenAI is.
    • gilrain 39 minutes ago
      Why are people so greedy? Why can't they wait to be given something, or remain content without, rather than take from others by force? I'm genuinely curious too.
      • diydsp 30 minutes ago
        Animalistic survival desires under conditions of engineered scarcity and reduced awareness. Compounded with powerful need for consistency - to keep everything this way.

        Anyone with a decent amount of social power is aware and skilled at maintaining these conditions whether consciously or not.

        The praxes are non-trivial: awareness, bravery, flexibility, and seflessness.

        This could be summarized as "unhealthy competition."

    • watwut 43 minutes ago
      They are releasing their work for the profit of the humanity. They do not like OpenAI claiming credit for it. After all, profit of the humanity and profit of the OpenAI are two different things. Maybe even mutually exclusive at this point.

      And lying about whether the proof was done by mathematicians or OpenAI is bad humanity, so.