LLMs Unreliable for Cosmetic Chemistry and Skincare Advice

Amelia Liu· August 18, 2026 View original

Key takeaways

  • LLMs perform poorly in cosmetic chemistry and skin health, especially in quantitative tasks.
  • Responses lack technical depth for informed consumer decisions.
  • Authoritative but incorrect outputs pose a significant risk to users.
  • General-purpose LLMs are currently unreliable for specialized skincare advice.

Who benefits

Consumer TechE-commerceCosmeticsHealthcarePharmaceuticals

Summary

A benchmarking study of 14 LLMs on cosmetic chemistry and skin health found overall poor performance, especially in quantitative reasoning and structural identification. While general skincare questions received reasonable answers, responses lacked technical depth, and authoritative-sounding errors posed a risk, making LLMs unreliable for informed consumer decisions.

With consumers increasingly seeking skincare advice from AI chatbots, a critical evaluation of Large Language Models (LLMs) in cosmetic chemistry and skin health was conducted. A study benchmarked 14 LLMs on structured topics, including chemical properties of ingredients and common skincare scenarios, with web search disabled to assess internalized knowledge. The overall performance was found to be poor, with significant weaknesses in quantitative reasoning and identifying chemical structures. Although LLMs could provide generally reasonable answers to broad skincare questions, their responses consistently lacked the technical depth necessary for consumers to make truly informed decisions. A notable risk identified was that authoritative-sounding but technically incorrect outputs were less likely to prompt skepticism than responses acknowledging uncertainty. The findings strongly suggest that general-purpose LLMs, primarily trained on unverified public data, are currently not dependable sources for cosmetic chemistry information. Future improvements will likely require fine-tuning on verified chemical and dermatological datasets, alongside substantial advancements in algorithmic reasoning.

Why it matters

Professionals in consumer tech, e-commerce, and healthcare must exercise extreme caution when considering LLMs for providing advice in specialized, sensitive domains like skincare. The risk of disseminating authoritative-sounding but incorrect information can harm consumers and damage brand trust.

How to implement this in your domain

  1. 1Implement strict content review and fact-checking for any LLM-generated advice in health or chemistry-related fields.
  2. 2Avoid deploying general-purpose LLMs for direct consumer advice in specialized domains without extensive fine-tuning on verified data.
  3. 3Develop specialized LLMs or hybrid systems that integrate expert knowledge bases for accurate technical information.
  4. 4Educate users about the limitations of AI-generated advice and encourage consultation with human experts.
  5. 5Prioritize transparency by clearly indicating when information is AI-generated and its potential for error.

Original post by Amelia Liu

"arXiv:2608.14631v1 Announce Type: new Abstract: As consumers increasingly turn to AI chatbots for skincare advice, the technical accuracy of Large Language Models (LLMs) in cosmetic chemistry remains largely under-evaluated. We benchmarked 14 LLMs on a structured set of topics re…"

View on X

Originally posted by Amelia Liu on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses