نابودی مکانیکی تاریخ: چگونه توسعهٔ هوش مصنوعی میراث مکتوب بشر و نیروی کار جهانی را میبلعد
از «اسکن مخرب» و خصوصیسازی حافظهٔ جمعی تا کار دو دلاری و امپریالیسم دادهمحور
رقابت افسارگسیختهٔ غولهای فناوری برای آموزش مدلهای بزرگ زبانی وارد مرحلهای شده است که در آن کتاب نه بهعنوان میراث فرهنگی، بلکه همچون مادهٔ خام یک خط تولید صنعتی دیده میشود. نسخهٔ چاپی خریداری میشود، عطف آن زیر تیغه میرود، صفحات به داده تبدیل میشوند و پیکر کتاب پس از پایان استخراج کنار گذاشته یا بازیافت میشود. در سوی دیگر همین زنجیره، انسانهایی در انبارها، مراکز اسکن و پلتفرمهای برچسبگذاری، کار مادی و روانی سنگینی انجام میدهند تا محصول نهایی به نام «هوش مصنوعی» و به مالکیت چند شرکت انحصاری عرضه شود.
بحث فقط بر سر کپیرایت نیست. با یک الگوی انباشت روبهرو هستیم: میراث جمعی بشر بهطور انبوه استخراج میشود؛ هزینهٔ آمادهسازی آن به کارگران و جوامع کمقدرت منتقل میگردد؛ و محصول نهایی در مدلهای بسته و سودآور متمرکز میشود. شواهد محکم دربارهٔ نابودی فیزیکی نسخهها و استفاده از کتابهای دزدی وجود دارد؛ اما مدرک عمومی کافی نداریم که نشان دهد شرکتها عمداً متن اصلی کتابهای مشخصی را پس از اسکن بازنویسی کردهاند. آنچه مستند است، خطا و حذف در دیجیتالسازی و بازتولید ناقص، افزوده و تکراری کتابها از سوی مدلهاست.
۱. پروژهٔ پاناما: کارخانهای برای تبدیل کتاب به توکن
در اوایل ۲۰۲۴، آنتروپیک پروژهای با نام رمزی Panama راه انداخت. یکی از اسناد داخلیِ علنیشده هدف آن را تلاش برای «اسکن مخرب همهٔ کتابهای جهان» توصیف کرد و بر محرمانهماندن پروژه تأکید داشت. خریدها از فروشندگان بزرگ کتاب دستدوم، از جمله Better World Books و World of Books، در محمولههای دههاهزارجلدی انجام شد.
منطق اقتصادی روشن بود: خرید یک نسخهٔ دستدوم و نابودکردن آن میتوانست از مذاکره برای مجوز میلیونها اثر ارزانتر و سریعتر باشد. پیمانکار عطف را جدا میکرد، صفحات را میبرید، آنها را اسکن و با OCR به داده تبدیل میکرد. پیشنهاد یکی از پیمانکاران پردازش پانصد هزار تا دو میلیون کتاب در شش ماه را در نظر داشت.
قاضی ویلیام آلساپ در ژوئن ۲۰۲۵ آموزش مدل با کتابهای قانونی خریداری و اسکنشده را در این پرونده «استفادهٔ منصفانه» دانست؛ اما نگهداری میلیونها نسخهٔ دریافتشده از کتابخانههای سایه مانند LibGen و PiLiMi با همین استدلال تطهیر نشد. پروندهٔ نسخههای دزدی به توافقی ۱٫۵ میلیارد دلاری دربارهٔ صدهاهزار عنوان انجامید.
۲. چرا نابودی یک نسخهٔ چاپی صرفاً نابودی کاغذ نیست
کتاب فقط ظرفی برای واژهها نیست. چاپ، نوبت ویرایش، صفحهآرایی، کاغذ، صحافی، حاشیهنویسی، مهر کتابخانه، امضا، اهدانامه و حتی لکهها میتوانند بخشی از تاریخ اجتماعی یک اثر باشند. OCR بیشتر این اطلاعات را حذف میکند.
برای سرمایهدار و شرکت هوش مصنوعی، کتاب در لحظهٔ خرید به «کالا» و سپس مادهٔ خام داده تبدیل میشود؛ ارزش آن با قیمت خرید، هزینهٔ اسکن و بازده آموزشی سنجیده میشود. اما از دید جامعهٔ بشری، همان کتاب حامل زبان، حافظه، تجربه، هنر و تاریخ نسلهاست. ارزش مبادلهای آن ممکن است ناچیز باشد، اما ارزش فرهنگی و تاریخیاش میتواند جایگزینناپذیر باشد. آنچه برای شرکت پس از استخراج داده زائد است، برای جامعه ممکن است سندی از زندگی و آگاهی انسان باشد.
«قدیمی»، «کمیاب» و «نسخهٔ یگانه» دقیقاً یک معنا ندارند؛ بااینحال کتابی با پنجاه، صد یا دویست سال قدمت کالای مصرفی معمولی نیست. در خریدهای میلیونجلدی و بدون ارزیابی عمومی نسخهبهنسخه، کتابهای قدیمی، خارج از چاپ، کمیاب در بازار و بالقوه یگانه همگی در معرض تیغه قرار میگیرند. نبود فهرست عمومی دقیق دفاع شرکت نیست؛ بخشی از اتهام است. پیش از برش، شرکت باید ثابت کند نسخه ارزش آرشیوی، قدمت ویژه، حاشیهنویسی تاریخی یا کمیابی ندارد.
۳. نیروی کار دو دلاری و کارخانهٔ پنهان هوش مصنوعی
پشت تصویر خودکار هوش مصنوعی، شبکهای مادی قرار دارد: خرید، حملونقل، انبار، موجودی، جداسازی عطف، برش، اسکن، کنترل کیفیت، OCR، پاکسازی، برچسبگذاری، پالایش محتوای سمی و بازیافت. بدون این دستها، «ابر» و «مدل» حتی یک صفحه را به داده تبدیل نمیکنند.
دستمزد کارکنان مشخص پروژهٔ پاناما علنی نشده است. اما تحقیق TIME نشان داد OpenAI از طریق Sama در کنیا کارگرانی را برای خواندن و برچسبزدن متنهای خشونتآمیز و آزاردهنده به کار گرفت که حدود ۱٫۳۲ تا ۲ دلار در ساعت دریافت میکردند و از آسیب روانی جدی خبر دادند. در فیلیپین نیز واشنگتنپست از هزاران کارگر Remotasks/Scale AI گزارش داد که با پرداختهای بهتعویقافتاده، کاهشیافته یا لغوشده و درآمدهایی گاه پایینتر از حداقل دستمزد مواجه بودند. این نمونهها نشان میدهند خطر و فرسودگی در پایین زنجیره و مالکیت، اعتبار و سود در بالای آن متمرکز است.
۴. آیا متن کتابها را تغییر دادهاند؟ سه پدیدهٔ متفاوت
الف) تغییر عمدی نسخهٔ اصلی
سند عمومی قابلاتکایی پیدا نشده که نشان دهد یک شرکت بزرگ پس از اسکن، متن یک عنوان مشخص را عمداً سانسور یا بازنویسی و جایگزین اصل کرده است. چنین ادعایی بدون عنوان، چاپ، صفحه و فایل مقایسهای اعتبار نقد را تضعیف میکند.
ب) خطا و حذف در دیجیتالسازی
این پدیده واقعی است. کتابخانهٔ کلمبیا نمونههایی از Google Books را گزارش کرده که صفحات تاخورده باز نشده و محتوای زیر آنها غایب مانده است. تصاویر ناقص، دست اپراتور، صفحههای کج و خطاهای OCR نیز مستند شدهاند. اگر اصل فیزیکی نابود شود، امکان بازبینی و اصلاح ضعیف میشود.
ج) بازسازی کتاب از حافظهٔ مدل
مدل زبانی آرشیو نیست. پژوهشی در ۲۰۲۶ توانست بخشهای بزرگی از دوازده کتاب را از Claude 3.7، GPT‑4.1، Gemini 2.5 Pro و Grok 3 استخراج کند. برای Harry Potter and the Sorcerer’s Stone، پژوهشگران ۹۵٫۸ درصد از کلود، ۷۶٫۸ درصد از جمینای، ۷۰٫۳ درصد از گروک و حدود ۴ درصد از GPT‑4.1 استخراج کردند. در The Great Gatsby، کلود پوشش ۹۷٫۵ درصدی داشت اما بخشهایی از صفحات ۱۱۴ تا ۱۳۲ را تکرار کرد و خروجی شامل افزودهها و قسمتهای جاافتاده بود. حافظهٔ مدل نسخهٔ اصیل نیست.
۵. آنچه نمونههای عینی ثابت میکنند
- پروژهٔ پاناما: خرید و اسکن مخرب میلیونها کتاب و دورریزی یا بازیافت نسخههای کاغذی مستند است؛ نایاببودن همهٔ نسخهها مستند نیست.
- کتابخانههای سایه: منشأ، رضایت و نحوهٔ تحصیل داده اهمیت حقوقی و اخلاقی دارد.
- The Great Gatsby: پوشش بالا همراه با تکرار، افزوده و حذف؛ خروجی مدل جای نسخهٔ مرجع نیست.
- Google Books و OCR: دیجیتالسازی بدون حفظ اصل و کنترل کیفیت میتواند دادهٔ ناقص را به مرجع غالب تبدیل کند.
۶. چرخهٔ امپریالیسم دادهمحور
در مرحلهٔ نخست، ثروت فرهنگی—کتاب، تصویر، زبان و تجربهٔ اجتماعی—از سراسر جهان استخراج میشود. در مرحلهٔ دوم، پاکسازی و برچسبگذاری به کارگران ارزانتر و کمقدرتتر واگذار میگردد. در مرحلهٔ سوم، محصول این تولید اجتماعی در قالب مدل و زیرساخت ابری به مالکیت چند شرکت عمدتاً مستقر در شمال جهانی درمیآید و دوباره به جامعه فروخته میشود.
جامعه طی قرنها کتاب را نوشته، ترجمه، چاپ، تدریس و نگهداری کرده و کارگران امروز آن را به داده تبدیل میکنند؛ اما فهرست منابع، مدل، زیرساخت و سود خصوصی میماند. شرکت نهفقط کتاب، بلکه شیوهٔ دیدن آن را کنترل میکند.
۷. پاسخهای رایج شرکتها و محدودیت آنها
مدافعان میگویند نسخههای خریداریشده فراوان و کمارزشاند، اسکن مخرب سابقه دارد و مالک میتواند نسخهٔ خود را نابود کند. اما عملی ممکن است قانونی و در مقیاس میلیونها نسخه از نظر فرهنگی نامسئولانه باشد. اگر نسخهها فراواناند، ثبت عمومی آن باید آسان باشد. اگر هدف حفظ دانش است، چرا تصاویر و فرادادهها به آرشیو عمومی سپرده نمیشوند؟ و اگر پروژه بیخطر است، چرا اسناد داخلی بر محرمانگی تأکید داشتند؟
۸. حداقل قواعد برای جلوگیری از فراموشی صنعتی
- اسکن غیرمخرب اصل باشد و نابودی فقط پس از ارزیابی مستقل و اثبات فراوانی همان چاپ مجاز شود.
- عنوان، ناشر، سال، چاپ، شابک، منشأ و وضعیت فیزیکی هر جلد در ثبت عمومی قرار گیرد.
- آثار نایاب، چاپهای محلی، نسخههای حاشیهنویسیشده و فاقد نسخهٔ آرشیوی از خط تخریب خارج شوند.
- تصویر صفحه، OCR، فراداده، هش و گزارش خطا به آرشیو عمومی مورد اعتماد سپرده شود.
- منشأ، چاپ و کیفیت OCR در مدل قابلردیابی باشد و حسابرسی مستقل انجام شود.
- حقوق نویسنده، ناشر و کارگر—شامل رضایت، جبران، ایمنی، دستمزد پایه و حق تشکل—تضمین شود.
- کتابداران، نویسندگان، مترجمان، پژوهشگران، اتحادیهها و جوامع زبانی در قواعد مشارکت کنند.
۹. نتیجهگیری: در برابر بلعیدن تاریخ و کار انسانی
نابودی کتاب برای تغذیهٔ مدل بسته تصمیمی صرفاً فنی نیست. تیغه شیرازه را میبُرد؛ OCR متن را از جسم و تاریخش جدا میکند؛ کارگر دو دلاری مواد سمی را پالایش میکند؛ و محصول نهایی با مالکیت شرکت انحصاری وارد بازار میشود. آنچه جمعی تولید شده، خصوصی تصاحب میشود.
در جامعهٔ سرمایهداری، سودآوری تنها هدف نیست، اما در صدر سلسلهمراتب تصمیمگیری قرار دارد. کتاب قدیمی میتواند برای شرکت کالایی چنددلاری و پس از اسکن کاغذ بیمصرف باشد؛ درحالیکه میلیاردها دلار برای بودجهٔ نظامی، جنگ، رقابت تسلیحاتی و حفاظت از مرز، خاک و پرچم هزینه میشود. میراث فرهنگی تا جایی مورد توجه قرار میگیرد که سودآور، قابلفروش یا ابزار قدرت باشد.
اما خاک بدون حافظه، پرچم بدون فرهنگ و مرز بدون انسانهای آگاه، پوستههایی تهیاند. آگاهی اجتماعی فقط انبارشدن اطلاعات نیست؛ رابطهٔ زندهٔ انسان با متن، زبان، تاریخ و شرایط مادی پدیدآمدن اثر است. این میراث را نمیتوان یکشبه با اسکن و تبدیل به توکن از نسلی به نسل دیگر منتقل کرد.
در «نظم اجتماعی نوین در عصر آگاهی»، آثار فرهنگی و دادههای انسانی سراسر جهان سرمایهٔ خصوصی چند شرکت نخواهند بود؛ بخشی از ثروت مشترک و حافظهٔ اجتماعی بشر خواهند بود. وظیفهٔ فناوری بلعیدن و انحصاریکردن میراث نیست؛ حفظ زمینه، گسترش دسترسی عمومی، پیوند تجربهها و توانمندکردن جامعه برای تصمیمگیری آگاهانه است.
راهحل دشمنی با فناوری نیست؛ شکستن رابطهای است که فناوری را به ماشین انباشت بدل میکند. کتاب پس از اسکن نباید زباله شود، کارگر نباید مصرفشدنی باشد و حافظهٔ مدل نباید جای نسخهٔ اصیل را بگیرد.
منابع منتخب
- Federal court order in Bartz v. Anthropic (June 2025)
- Washington Post investigation of Project Panama
- Euronews report on unsealed Project Panama documents
- Authors Guild guide to the Anthropic case and settlement
- Associated Press report on fair use and pirated copies
- Ahmed et al., Extracting Books from Production Language Models (2026)
- Columbia University Libraries on incomplete Google Books
- Research on OCR errors in HathiTrust and Project Gutenberg
- TIME investigation of OpenAI/Sama workers in Kenya
- Washington Post report on Remotasks/Scale AI workers in the Philippines
Mechanical Destruction of History: How AI Development Is Devouring Humanity’s Written Heritage and Global Labor
From destructive scanning and the privatization of collective memory to two-dollar labor and data imperialism
The unrestrained race among technology giants to train large language models has entered a disturbing phase. Books are no longer treated primarily as cultural heritage, but as raw material for an industrial production line. A printed copy is purchased, its spine is placed under a blade, its pages are converted into data, and its physical body is discarded or recycled after extraction. Elsewhere in the same chain, people in warehouses, scanning centers, and data-labeling platforms perform exhausting physical and psychological work so the finished product can be marketed as “artificial intelligence” and owned by a handful of corporations.
This is not merely a copyright dispute. It is a pattern of accumulation: humanity’s collective heritage is extracted at scale; the cost of preparing it is shifted onto workers and less powerful communities; and the final product is concentrated in closed, profitable models. Strong evidence documents the physical destruction of books and the acquisition of pirated copies. There is not, however, sufficient public evidence that major companies intentionally rewrote the original text of identified books after scanning. What is documented is error and omission in digitization, and incomplete, repetitive, or augmented reconstruction by language models.
1. Project Panama: A Factory for Turning Books into Tokens
In early 2024, Anthropic launched a project code-named Panama. An internal document later made public described its aim in startling terms: an effort to “destructively scan all the books in the world.” The documents also emphasized secrecy. Books were acquired in shipments of tens of thousands from major secondhand sellers, including Better World Books and World of Books.
The economic logic was straightforward. Buying and destroying a used copy could be faster and cheaper than negotiating licenses for millions of works. A contractor removed the binding, cut the pages, scanned them, and converted the images into text through optical character recognition. One vendor proposal contemplated processing roughly 500,000 to two million books in six months. Court records stated that Anthropic bought and scanned millions of printed books and retained searchable digital copies for its central library.
The legal distinction matters. In June 2025, Judge William Alsup found that training on lawfully purchased and scanned books qualified as fair use in the case before him. But the retention of millions of copies previously obtained from shadow libraries such as LibGen and PiLiMi was not cleansed by the same argument. Litigation over pirated copies ultimately led to a $1.5 billion settlement involving hundreds of thousands of titles. “Buy and scan” and “download a pirated copy” were legally different, although both belonged to the same race to accumulate data.
2. Destroying a Printed Copy Is Not Merely Destroying Paper
A book is not simply a container for words. Edition, layout, paper, binding, marginal notes, library stamps, signatures, inscriptions, price labels, and even stains can form part of a work’s social history. OCR text removes most of this information, and even page images do not necessarily preserve a book’s physical construction, folded materials, insertions, or evidence of ownership and use.
For the capitalist investor and the AI company, the book becomes a commodity at purchase and then raw material for data. Its value is measured by purchase price, scanning cost, and training yield. For human society, however, the same book carries language, memory, experience, art, and the history of generations—values that cannot be reduced to a few dollars on the secondhand market. Its exchange value may be tiny while its cultural and historical value is irreplaceable. What becomes waste to the corporation after data extraction may remain evidence of human life and consciousness to society.
“Old,” “scarce,” and “unique copy” are not identical bibliographic categories; age alone does not establish how many copies survive. Nevertheless, a book that is fifty, one hundred, or two hundred years old is not an ordinary disposable commodity. In million-volume acquisitions without public copy-by-copy review, old, out-of-print, market-scarce, and potentially unique books all pass under the same blade.
The unsealed record does not provide a complete public inventory of title, edition, year, physical condition, and the number of surviving copies. We therefore cannot state precisely how many rare copies were destroyed. But ignorance is not a defense; it is part of the indictment. A corporation destroying written history on an industrial scale should bear the burden of proving before cutting that a copy lacks archival value, unusual age, historically significant annotations, or scarcity.
There is also a loss of public circulation. A secondhand book could have been resold, donated, or placed in a library. After destructive scanning, the text becomes a private corporate asset while the physical book is no longer usable. This is not preservation unless high-quality images, exact edition metadata, error reports, and durable public access are also secured. A training pipeline is designed to produce tokens; an archive is designed to preserve an identifiable work and its context.
3. Two-Dollar Labor and AI’s Hidden Factory
The promotional image of AI is an autonomous algorithm. The reality is a material network of purchasing, loading, transportation, warehouse rent, inventory, debinding, cutting, scanning, quality control, OCR, file cleaning, labeling, toxic-content filtering, and recycling. Project Panama records confirm logistics management, warehouse labor, and scanning contractors. Without these hands, “the cloud” cannot turn a single paper page into data.
The wages of Project Panama’s own workers have not been publicly disclosed, so the two-dollar figure should not be attributed directly to them. But another part of the same industry is well documented. A TIME investigation found that OpenAI, through its contractor Sama, used workers in Kenya to read and label text involving violence, hate, and sexual abuse. Their take-home pay was approximately $1.32 to $2 per hour, and workers reported serious psychological harm. OpenAI paid the contractor more per hour, but only a small portion reached the worker. The gap condenses the architecture of AI accumulation: risk and exhaustion at the bottom; ownership, prestige, and profit at the top.
In the Philippines, the Washington Post reported that thousands of workers on Scale AI’s Remotasks platform faced delayed, reduced, or canceled payments and earnings that sometimes fell below local minimum wages. These workers did not necessarily correct Panama’s scanned books. Their experience nevertheless shows how the global data-production line exploits wage differences, weak legal protection, and outsourced responsibility.
4. Did Companies “Change” the Books? Three Different Phenomena
A. Intentional alteration of an original
No reliable public evidence identified during this research shows that Anthropic or another major company deliberately censored or rewrote the text of a named book after scanning and retained the altered version as the original. Making that allegation without a title, edition, page, and comparative file weakens the broader critique.
B. Error and omission in digitization
This phenomenon is real and predates generative AI. Columbia University Libraries documented Google Books copies in which folded pages were not opened, leaving the material beneath them absent. Other projects have documented operators’ hands covering text, skewed and incomplete pages, and OCR errors. Research on large HathiTrust and Project Gutenberg corpora has also found character, word, and duplication errors. These are not necessarily intentional falsifications, but if the physical original has been destroyed, correction becomes much harder.
C. Reconstructing a book from model memory
A language model is not an archive. A 2026 study extracted substantial portions of twelve books from Claude 3.7 Sonnet, GPT‑4.1, Gemini 2.5 Pro, and Grok 3. For Harry Potter and the Sorcerer’s Stone, researchers extracted 95.8% from Claude, 76.8% from Gemini, 70.3% from Grok, and about 4% from GPT‑4.1. Even high recovery did not produce a pristine edition. Claude’s reconstruction of The Great Gatsby reached 97.5% by the study’s metric but extensively repeated material corresponding to pages 114–132 and contained both additions and omissions. Model memory is not an authentic copy.
5. What the Concrete Examples Actually Establish
- Project Panama: The purchase and destructive scanning of millions of books, retention of files, and disposal or recycling of paper copies are documented. It is not documented that every copy was rare.
- Shadow libraries: Source, consent, and acquisition method matter legally and ethically.
- The Great Gatsby: High recovery accompanied by repetition, additions, and omissions demonstrates that model output is not a reference edition.
- Google Books and OCR: Digitization without preservation and quality control can make incomplete data the dominant reference.
6. The Cycle of Data Imperialism
This chain can be described as a new form of data imperialism. First, cultural wealth—books, images, languages, and social experience—is extracted from around the world. Second, cleaning, labeling, and safety work is shifted to cheaper and less powerful workers. Third, the result of this social production becomes the property of a small number of corporations, mostly headquartered in the Global North, and is sold back to society through proprietary models, interfaces, and cloud infrastructure.
Society spent centuries writing, translating, printing, teaching, criticizing, and preserving books; today’s workers convert them into usable data. Yet the source list, model, infrastructure, and profit remain private. The company controls not only access to the book but also how it is seen: which edition enters the corpus, what is removed as a duplicate, which signs OCR ignores, and under what rules the model responds. Even without a censorship conspiracy, concentrating such decisions in opaque institutions creates immense epistemic and economic power.
7. Common Corporate Defenses—and Their Limits
Defenders argue that acquired books were generally plentiful, low-value secondhand copies; destructive scanning has precedents; owners may destroy their property; and a model learns patterns rather than acting as a public library. The court accepted part of the transformative-use argument in Anthropic’s case.
These defenses do not settle the public-policy question. Conduct may be legal in one domain and culturally irresponsible at the scale of millions of copies. If the copies were truly plentiful, a public registry should be easy. If preservation was the goal, why were page images and metadata not deposited with a trusted public archive? If the model only “learns,” why could researchers extract near-complete copyrighted books? And if the project was harmless, why did internal documents stress secrecy?
8. Minimum Rules Against Industrial Forgetting
- Non-destructive scanning must be the default; destruction should require independent assessment and proof that the exact edition is abundant.
- Title, author, publisher, year, edition, ISBN, acquisition source, and physical condition of each copy should be publicly registered.
- Rare works, local editions, annotated or signed copies, and books without archival copies must be removed from destructive lines.
- Page images, OCR, metadata, file hashes, and error reports should be deposited with a trusted public archive, subject to appropriate copyright access rules.
- Models should preserve source and edition provenance, OCR quality, and a distinction between quotation and probabilistic reconstruction.
- Independent audits should compare digital files with physical originals and publish missing-page and OCR-error rates.
- Authors and publishers need consent, opt-out, compensation, and collective licensing mechanisms.
- Workers need contractor transparency, safety standards, base wages, workload limits, mental-health protection, and organizing rights.
- Librarians, writers, translators, researchers, labor unions, and language communities must participate in collection and destruction rules.
9. Conclusion: Against the Devouring of History and Human Labor
Destroying physical books to feed closed models is not a merely technical decision. The blade cuts the spine; OCR separates text from body and history; the two-dollar worker filters toxic material; and the finished product enters the market under corporate ownership. What was socially produced is privately appropriated.
In capitalist society, profitability is not the only objective, but it sits at the top of the decision-making hierarchy. An old book may become a few-dollar commodity and, after scanning, useless paper, while the same society spends billions on military budgets, war, arms competition, and the defense of borders, soil, and flags. Humanity’s cultural heritage—books, languages, music, paintings, stories, local memory, and generational experience—is valued to the extent that it is profitable, marketable, or useful to power. What offers no immediate return is exposed to neglect, privatization, or destruction.
But soil without memory, a flag without culture, and borders without conscious human beings are empty shells. Humanity’s cultural and artistic achievements are the result of centuries of life, suffering, struggle, imagination, and cooperation. They cannot be transmitted between generations overnight by scanning them, reducing them to tokens, and storing them in a virtual model. Social consciousness is not an information warehouse; it is humanity’s living relationship with text, language, history, experience, and the material conditions that produced them.
In A New Social Order in the Age of Consciousness, cultural works and human data scattered across the world would not be the private capital of a few corporations. They would be treated as common wealth and part of humanity’s social memory. A local edition, a less widely spoken language, a reader’s annotation, the story of a small community, or a worker’s experience in a distant place may yield little profit, yet carry enormous value for the growth of social intelligence. Technology’s task would not be to devour and monopolize this heritage, but to preserve context, expand public access, connect human experiences, and empower society to understand and decide consciously.
The solution is not hostility to technology or the suspension of knowledge. It is to break the social relationship that turns technology into a machine of accumulation. Books must not become waste after scanning; workers must not become disposable; and model memory must never replace authentic, verifiable editions.
Selected Sources
- Federal court order in Bartz v. Anthropic (June 2025)
- Washington Post investigation of Project Panama
- Euronews report on unsealed Project Panama documents
- Authors Guild guide to the Anthropic case and settlement
- Associated Press report on fair use and pirated copies
- Ahmed et al., Extracting Books from Production Language Models (2026)
- Columbia University Libraries on incomplete Google Books
- Research on OCR errors in HathiTrust and Project Gutenberg
- TIME investigation of OpenAI/Sama workers in Kenya
- Washington Post report on Remotasks/Scale AI workers in the Philippines
