AI Companies Destroy Books; Scan Rare Collections Before It's Too Late
Key takeaways
- AI data acquisition can lead to the destruction of physical artifacts.
- Rare books are particularly vulnerable to damage during scanning.
- Urgent action is needed to digitize and preserve unique collections.
- Ethical data sourcing practices are crucial for AI companies.
Who benefits
Summary
AI companies are reportedly destroying physical books during the process of scanning them for training data, raising concerns about the loss of rare and unique collections. There's an urgent call to digitize these valuable books before they are potentially lost forever.
Why it matters
Professionals in cultural heritage, libraries, and even AI development should care about ethical data sourcing and the preservation of irreplaceable knowledge. This highlights the collateral damage of unchecked data acquisition.
How to implement this in your domain
- 1Audit current data acquisition practices to ensure ethical handling of physical materials.
- 2Collaborate with libraries and archives to fund and execute high-quality digitization projects.
- 3Develop and implement best practices for handling physical media during scanning processes.
- 4Investigate alternative data sourcing methods that minimize physical impact.
Original post by Cider9986
"AI companies destroy physical books – let's scan rare books before it's too late"
View on XOriginally posted by Cider9986 on X · view source
Want to go deeper?
Turn these trends into skills with Learnijoy's hands-on AI & tech courses.
Explore coursesMore in AI News & Tools
New York Study Finds Flaws in AI-Based Lead Pipe Classification
A study in New York State audited predictive models used by utilities to classify lead service lines, finding significant discrepancies where models contradicted physical verification, especially in New York City. The research highlights that many addresses classified by models as "Known Other" or without lead were in older buildings where lead is expected.
Language Models Leak Sensitive Data from Context Window.
Research reveals that large language models can inadvertently leak sensitive user data present in their context window, even when explicitly refusing direct extraction. Adversaries can exploit this leakage through novel adaptive attacks, reconstructing secrets from seemingly benign outputs.
FleetSieve Optimizes LLM Fleet Configuration with SLO-Aware Profiling.
FleetSieve is a new profiling method that efficiently configures LLM serving fleets by selectively measuring performance based on its expected impact on resource allocation and Service Level Objectives (SLOs). It significantly reduces profiling time compared to exhaustive or random methods while ensuring SLO compliance and maximizing throughput.