// wiki · analysis · September 2026
Training a model on someone else's works
Draft article. The description of the norms of the law on supporting AI follows public reports of summer 2026; it is verified against the official text before publication.
A draft for discussion. The article's theses are a map of the question, not a conclusion about the legality of a specific dataset: the assessment always depends on the composition of the data and the chain of its acquisition.
The short answer
Training a model on someone else's work is a use of it, and by the general rule the rightsholder's consent is needed. There is no general “for machine learning” exception in the Civil Code. The 2026 law on supporting AI technologies softened the situation, but narrowly: without consent one may analyse legally obtained materials from open access — and only to train the (“sovereign”) systems defined in the law. For other cases the dataset is assembled legally: licences, purchased data, synthetic data.
What the Civil Code says today
Article 1273 of the Civil Code permits free reproduction of a work for the personal needs of a citizen user — training a commercial model cannot be fitted into that norm: the purpose is not personal, and the user is usually not the one training the model. Article 1274 lists the cases of free use (quotation, parody, educational purposes and others) — training AI systems is not in the list. Everything outside the list requires the rightsholder's consent (a licence).
What the law on supporting AI changes
According to reports of summer 2026, the law on supporting the development of AI technologies (bill No. 1271570-8) allows developers to work with others' works without rightsholders' consent — extracting data, comparing and analysing the material — provided three conditions are met simultaneously:
- the copy of the work was obtained legally;
- the material is in open access;
- the material is needed only for training sovereign AI systems (public discussion cited large Russian models as examples).
What follows practically: first, the exception is not universal — a product built on a foreign model, or a model for a narrow internal service, may not fall under the wording about “sovereign systems”; second, permission to analyse is not the same as the right to store copies and reproduce the dataset. While the article of the law has not been applied in public practice, the conservative approach is to assemble the dataset as if the exception did not exist.
Data is not only works
Not everything in a dataset is a work, and non-works have their own regimes:
- Personal data — 152-FZ: purposes and grounds of processing, and when it gets into a model — questions of deletion and localisation.
- Databases — the maker of a database holds a related right (Articles 1333–1334 of the Civil Code): extracting a substantial part of someone else's database infringes that right even where the individual facts are not protected.
- Secrets — a dataset of internal company data is subject to the trade secret regime and NDAs with partners.
Risks
- A rightsholder's claim — compensation without proof of damages from 10 thousand to 5 million roubles per object (Article 1301 of the Civil Code); with a large dataset the arithmetic works against you.
- Stopping the product — a court may prohibit the use of a model trained on disputed data; retraining costs more than the dispute.
- A problem in a deal — in an investment round or a company sale, a dataset without a provenance history becomes a due diligence red flag: cf. “Due diligence of an AI asset”.
A data provenance log
The working risk-management instrument is a log for each dataset: source and method of acquisition, licence or consent (with requisites), use restrictions, extraction date, volume. The same log answers the questions of a counterparty's lawyers and of the regulator. It is complemented by: licences from rightsholders and platforms, purchasing ready datasets with a verified licence, synthetic data where permissible, and filtering personal data before it enters the sample.
Checking a dataset and a model's training conditions is the task of technology rights practice; the team defines the scope of the check for a specific product after an enquiry.