| 1 |
Barocas, S. and Selbst, A. D. (2016). Big data’s disparate impact. California Law Review, 104(3):671–732.
|
|
| 2 |
Goodfellow, I., Bengio, Y., and Courville, A. (2016). Deep learning. MIT Press, Cambridge, MA, USA.
|
|
| 3 |
Chang, Y., Wang, X., Wang, J., Wu, Y., Yang, L., Zhu, K., Chen, H., Yi, X., Wang, C., Wang, Y., et al. (2024). A survey on evaluation of large language models. ACM Transactions on Intelligent Systems and Technology, 15(3):1–45.
|
|
| 4 |
Fehring, L., Frings, J., Rust, P., Kempny, C., Thürmann, P. A., and Meister, S. (2025). Extension of the consolidated criteria for reporting qualitative research guideline to large language models (COREQ+ LLM): Protocol for a multiphase study. JMIR Research Protocols, 14.
|
|
| 5 |
Guo, C., Pleiss, G., Sun, Y., and Weinberger, K. Q. (2017). On calibration of modern neural networks. In International Conference on Machine Learning, pages 1321–1330. PMLR.
|
|
| 6 |
Xu, Q., Soto, C., Shahnawaz, M., Liu, X., Jiang, X., and Kim, Y. (2025). Multi agent large language models for biomedical hypothesis generation in drug combination discovery. iScience, 28(12):113984.
|
|
| 7 |
Lipton, Z. C. and Steinhardt, J. (2019). Troubling trends in machine learning scholarship: Some ML papers suffer from flaws that could mislead the public and stymie future research. Queue, 17(1):45–77.
|
|
| 8 |
Madry, A., Makelov, A., Schmidt, L., Tsipras, D., and Vladu, A. (2018). Towards deep learning models resistant to adversarial attacks. In International Conference on Learning Representations.
|
|
| 9 |
McIntosh, T. R., Susnjak, T., Arachchilage, N. A. G., Liu, T., Xu, D., Watters, P., and Halgamuge, M. N. (2025). Inadequacies of large language model benchmarks in the era of generative artificial intelligence. IEEE Transactions on Artificial Intelligence.
|
|
| 10 |
Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K., and Galstyan, A. (2021). A survey on bias and fairness in machine learning. ACM Computing Surveys, 54(6):1–35.
|
|
| 11 |
Meng, K., Bau, D., Andonian, A., and Belinkov, Y. (2022). Locating and editing factual associations in GPT. Advances in Neural Information Processing Systems, 35:17359–17372.
|
|
| 12 |
Ovadia, Y., Fertig, E., Ren, J., et al. (2019). Can you trust your model’s uncertainty? Evaluating predictive uncertainty under dataset shift. Advances in Neural Information Processing Systems, 32.
|
|
| 13 |
Shi, W., Ajith, A., Xia, M., Huang, Y., Liu, D., Blevins, T., Chen, D., and Zettlemoyer, L. (2024). Detecting pretraining data from large language models. In International Conference on Learning Representations.
|
|
| 14 |
Springer, J. M., Goyal, S., Wen, K., Kumar, T., Yue, X., Malladi, S., Neubig, G., and Raghunathan, A. (2025). Overtrained language models are harder to fine-tune. arXiv preprint arXiv:2503.19206.
|
|
| 15 |
Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., et al. (2023). Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. Transactions on Machine Learning Research.
|
|
| 16 |
Tan, H., Zhan, S., Jia, F., Zheng, H.-T., and Chan, W. K. (2026). A hierarchical framework for measuring scientific paper innovation via large language models. Information Sciences, 728:122787.
|
|
| 17 |
Wang, H., Fu, T., Du, Y., Gao, W., Huang, K., Liu, Z., Chandak, P., Liu, S., Van Katwyk, P., Deac, A., et al. (2023). Scientific discovery in the age of artificial intelligence. Nature, 620(7972):47–60.
|
|
| 18 |
Weiss, K., Khoshgoftaar, T. M., and Wang, D. (2016). A survey of transfer learning. Journal of Big Data, 3(1):9.
|
|
| 19 |
Wiggins, W. F. and Tejani, A. S. (2022). On the opportunities and risks of foundation models for natural language processing in radiology. Radiology: Artificial Intelligence, 4(4).
|
|
| 20 |
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T., Cao, Y., and Narasimhan, K. (2023). Tree of thoughts: Deliberate problem solving with large language models. Advances in Neural Information Processing Systems, 36:11809–11822.
|
|
| 21 |
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., and Cao, Y. (2023). ReAct: Synergizing reasoning and acting in language models. In International Conference on Learning Representations.
|
|