What is this project about?
This project tests whether small datasets can be used to train CNN and AI models effectively. I use image augmentation to see which level of image variation affects the final result. To show that this can be done on my own principles, the project uses only open data from NASA and ESO, so no one's work is used without permission.
The problem
AI models are often trained on large datasets to improve performance and generalization. This practice
leads to work by human artists and writers being stolen by corporations. This project tries to address
this issue in two ways.
1. I only use open, publicly available NASA and ESO images.
2. I use augmentation (changing the existing images to create more variety) to make a small dataset
generate a larger amount of useful data, then test the results against benchmarks to see how the final
result is affected.
Idea
Image augmentation uses pre-existing data to create new data samples that can improve model optimization and generalizability.