I bought a datacenter GPU that doesn't fit in a normal motherboard, macgyvered the fan with jumper wires, and now I'm running a model that ties with Claude Sonnet 4.6 on benchmarks, all for £200.
You might not appreciate it, but if it’s posted online then it’s no different from someone else learning to code from reading the project. It’s not making copies of the code, it’s just strengthening the weights on a neural network. Sure, if the code is so obscure that nothing else is like it then it’s possible to get the model to regurgitate some of it due to having so few relevant sources, but it’s very unlikely to be comprehensive enough that it’s violating any copyright. If a court finds that to be the case, for some fictitious example, then I’m certain they can find an agreeable resolution to the isolated case.
However, none of that is justification for just writing off the technology entirely. Pandora’s Box has been opened. The genie isn’t going back in the bottle. You can’t close the barn door, all the cows already escaped. What do you think boycotting it will accomplish? What exactly is the goal by figuratively sticking your fingers in your ears and pretending the models don’t exist?
I have seen this argument before and it doesn’t make sense even in theory.
I used to work at a company that did open source work and also proprietary work with third party closed source code. The company didn’t let anyone who had seen proprietary code contribute to open source, because they felt once a person ‘learned’ from a proprietary codebase, then it’s too risky if similar looking code lands in a project.
Imagine if someone saw the source code for Excel. Then sometime later they notice that Calc didn’t have a feature that Excel did, and contributed an implementation of the feature. Even if they hadn’t been looking directly at the Excel source code in the moment of implementation, you think Microsoft would be so “understanding” when they see someone that once worked on Excel contributing what could be construed as infringing?
The AI companies also seem to acknowledge this, as they have offerings that promise not to use your proprietary code as training fodder. If it is not a risk of infringement, then why would it matter to promise that the proprietary code is kept out of training data? Though it was short lived, why would OpenAI have even made a deal to license Disney material if it’s all fair use anyway? If this sort of stuff is fair game, why do they get so pissy when other companies distill models?
Even as the AI company’s have roughly defended this scenario, their defense should be a cause for concern for users. Generally they say that anything they do with things they can read is ‘fair use’, and when exhibits of clearly infringing outputs are given, they respond with the model only did that because the user’s prompt directed it, and thus the responsibility for infringement should be with the AI user, not the engine that produced the infringing output. So the possibility of an unwitting infringement is possible as the AI companies explicitly say it’s the fault of the user even if it happens.
But we come to your last point, that essentially at this point, the whole thing is ‘too big to fail’ and thus the practical risk is low. Which is true. It’s just a bit disheartening that these companies are given free reign to interpret intellectual property law whichever way is convenient in the moment.
Thank you for taking the time to write a thoughtful, sincere response. I can tell you’ve given this some thought, and appreciate the fact you aren’t just regurgitating talking points.
You make valid points regarding copyright law concerns, but my own perspective is that it’s an even playing field now. No one individual, group, company, etc was targeted, or unduly affected relative to any other. It could be argued it was ethically wrong to have been done at all, but since it was, and it was done to EVERYONE, then in my view it is a shared creation. Everyone is equally entitled to the resulting, from the social media users whose conversations trained the models, up through the senior engineers or CEOs of non-profit organizations whose code trained the models on syntax.
None of it can be extracted reliably, none of it can be distilled - it is an amalgamation of information. Belonging you everyone. Saying you object to having unwillingly participated is understandable but meaningless since it cannot be undone, cannot be excised, and even if it could, the sheer amount on data means the elimination of any one source would hang an insignificant impact as to be noticable. So it’s moot. You may as well say you object to the Moon being named Luna in the past. Go for it, but it doesn’t change anything. Even if you convinced everyone it should have been called ‘Billy’ instead, you won’t change the reality of the past. Know what I mean?
I don’t mean to be disheartening by highlighting the futility of objection, but it’s undeniable. There’s zero benefit possible, zero gain, and zero impact. Why wallow in complaints of the indelible? Instead, accept and adapt, as humans excel at.
I’d personally like to see global tax laws that establish taxing of usage by businesses with that revenue used to fund Universal Basic Income for every person, so that the productivity and advances of this shared creation are fairly shared with everyone. I think that should be the goal of all objectors, because that is feasible, realistic and fair. Anything else is literally unrealistic. Like demanding of reality that gravity should give you special treatment. Such demands, while grand, are nonsensical.
You might not appreciate it, but if it’s posted online then it’s no different from someone else learning to code from reading the project. It’s not making copies of the code, it’s just strengthening the weights on a neural network. Sure, if the code is so obscure that nothing else is like it then it’s possible to get the model to regurgitate some of it due to having so few relevant sources, but it’s very unlikely to be comprehensive enough that it’s violating any copyright. If a court finds that to be the case, for some fictitious example, then I’m certain they can find an agreeable resolution to the isolated case.
However, none of that is justification for just writing off the technology entirely. Pandora’s Box has been opened. The genie isn’t going back in the bottle. You can’t close the barn door, all the cows already escaped. What do you think boycotting it will accomplish? What exactly is the goal by figuratively sticking your fingers in your ears and pretending the models don’t exist?
I have seen this argument before and it doesn’t make sense even in theory.
I used to work at a company that did open source work and also proprietary work with third party closed source code. The company didn’t let anyone who had seen proprietary code contribute to open source, because they felt once a person ‘learned’ from a proprietary codebase, then it’s too risky if similar looking code lands in a project.
Imagine if someone saw the source code for Excel. Then sometime later they notice that Calc didn’t have a feature that Excel did, and contributed an implementation of the feature. Even if they hadn’t been looking directly at the Excel source code in the moment of implementation, you think Microsoft would be so “understanding” when they see someone that once worked on Excel contributing what could be construed as infringing?
The AI companies also seem to acknowledge this, as they have offerings that promise not to use your proprietary code as training fodder. If it is not a risk of infringement, then why would it matter to promise that the proprietary code is kept out of training data? Though it was short lived, why would OpenAI have even made a deal to license Disney material if it’s all fair use anyway? If this sort of stuff is fair game, why do they get so pissy when other companies distill models?
Even as the AI company’s have roughly defended this scenario, their defense should be a cause for concern for users. Generally they say that anything they do with things they can read is ‘fair use’, and when exhibits of clearly infringing outputs are given, they respond with the model only did that because the user’s prompt directed it, and thus the responsibility for infringement should be with the AI user, not the engine that produced the infringing output. So the possibility of an unwitting infringement is possible as the AI companies explicitly say it’s the fault of the user even if it happens.
But we come to your last point, that essentially at this point, the whole thing is ‘too big to fail’ and thus the practical risk is low. Which is true. It’s just a bit disheartening that these companies are given free reign to interpret intellectual property law whichever way is convenient in the moment.
Thank you for taking the time to write a thoughtful, sincere response. I can tell you’ve given this some thought, and appreciate the fact you aren’t just regurgitating talking points.
You make valid points regarding copyright law concerns, but my own perspective is that it’s an even playing field now. No one individual, group, company, etc was targeted, or unduly affected relative to any other. It could be argued it was ethically wrong to have been done at all, but since it was, and it was done to EVERYONE, then in my view it is a shared creation. Everyone is equally entitled to the resulting, from the social media users whose conversations trained the models, up through the senior engineers or CEOs of non-profit organizations whose code trained the models on syntax.
None of it can be extracted reliably, none of it can be distilled - it is an amalgamation of information. Belonging you everyone. Saying you object to having unwillingly participated is understandable but meaningless since it cannot be undone, cannot be excised, and even if it could, the sheer amount on data means the elimination of any one source would hang an insignificant impact as to be noticable. So it’s moot. You may as well say you object to the Moon being named Luna in the past. Go for it, but it doesn’t change anything. Even if you convinced everyone it should have been called ‘Billy’ instead, you won’t change the reality of the past. Know what I mean?
I don’t mean to be disheartening by highlighting the futility of objection, but it’s undeniable. There’s zero benefit possible, zero gain, and zero impact. Why wallow in complaints of the indelible? Instead, accept and adapt, as humans excel at.
I’d personally like to see global tax laws that establish taxing of usage by businesses with that revenue used to fund Universal Basic Income for every person, so that the productivity and advances of this shared creation are fairly shared with everyone. I think that should be the goal of all objectors, because that is feasible, realistic and fair. Anything else is literally unrealistic. Like demanding of reality that gravity should give you special treatment. Such demands, while grand, are nonsensical.