Black Box to White Box: Discover Model Characteristics Based on Strategic Probing

In Machine Learning, White Box Adversarial Attacks rely on knowing underlying\nknowledge about the model attributes. This works focuses on discovering to\ndistrinct pieces of model information: the underlying architecture and primary\ntraining dataset. With the process in this paper, a structured set of input\nprobes and the output of the model become the training data for a deep\nclassifier. Two subdomains in Machine Learning are explored: image based\nclassifiers and text transformers with GPT-2. With image classification, the\nfocus is on exploring commonly deployed architectures and datasets available in\npopular public libraries. Using a single transformer architecture with multiple\nlevels of parameters, text generation is explored by fine tuning off different\ndatasets. Each dataset explored in image and text are distinguishable from one\nanother. Diversity in text transformer outputs implies further research is\nneeded to successfully classify architecture attribution in text domain.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC