When training data changes the story
A literary model study separates style corruption from factual knowledge poisoning, using GPT-2 and Llama.
Inside the project
Research showcaseOriginal artifact ↗
How it works.
Define the attack
Modify literary training material or factual question-answer data.
From the original project.
Saved artifacts · click to inspect
Original artifact ↗
Original artifact ↗
Original artifact ↗
Continue in the source.
Open the source notebook in Jupyter, Colab, or the environment described in the README. Data and model downloads may be required.
git clone https://github.com/eforus-overseer/Literary-LLM-Knowledge-Data-Poisoning.gitRead the setup and requirements ↗Project artifacts.
NLP_FinalProject_Efi_Adi_V2.ipynb ↗LLM_Data_Poisoning_QnA_Attack_Analysis.ipynb ↗Kafka_GPT_.ipynb ↗Data Poisoning Attacks on Literary Language Models.pdf ↗docs/Data_Poisoning_Attacks_on_Literary_Language_Models.pdf ↗Tolkien-GPT-Architecture.html ↗
Source links point to the original public repository. Credit belongs to the project authors and the dependencies credited there.