| 1 |
Chandra Sekhar Bhagavatula, Thanapon Noraset, and Doug Downey. 2013. Methods for exploring and mining tables on Wikipedia. In Proceedings of the ACM SIGKDD Workshop on Interactive Data Exploration and Analytics (IDEA '13). Association for Computing Machinery, New York, NY, USA, 18–26. https://doi.org/10.1145/2501511.2501516
|
|
| 2 |
Tianji Cong, Madelon Hulsebos, Zhenjie Sun, Paul Groth, and H. V. Jagadish. 2023. Observatory: Characterizing Embeddings of Relational Tables. Proc. VLDB Endow. 17, 4 (December 2023), 849–862. https://doi.org/10.14778/3636218.3636237
|
|
| 3 |
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
|
|
| 4 |
Thibault Formal, Benjamin Piwowarski, and Stéphane Clinchant. 2021. SPLADE: Sparse Lexical and Expansion Model for First Stage Ranking. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR '21). Association for Computing Machinery, New York, NY, USA, 2288–2292. https://doi.org/10.1145/3404835.3463098
|
|
| 5 |
Dayne Freitag, John Cadigan, Robert Sasseen, and Paul Kalmar. 2022. Valet: Rule-Based Information Extraction for Rapid Deployment. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pages 524–533, Marseille, France. European Language Resources Association.
|
|
| 6 |
Madelon Hulsebos. 2024. Table Representation Learning. Ph.D. Dissertation. University of Amsterdam, Amsterdam, The Netherlands. https://pure.uva.nl/ws/files/155198398/Thesis.pdf
|
|
| 7 |
Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2024. Lost in the Middle: How Language Models Use Long Contexts. Transactions of the Association for Computational Linguistics, 12:157–173.
|
|
| 8 |
Manning, C. D., Raghavan, P., and Schütze, H. (2008). Introduction to Information Retrieval. Cambridge University Press, Cambridge.
|
|
| 9 |
Norman W. Paton, Jiaoyan Chen, and Zhenyu Wu. 2023. Dataset Discovery and Exploration: A Survey. ACM Comput. Surv. 56, 4, Article 102 (April 2024), 37 pages. https://doi.org/10.1145/3626521
|
|
| 10 |
Mushtari Sadia, Zhenning Yang, Yunming Xiao, Ang Chen, and Amrita Roy Chowdhury. 2025. SQUiD: Synthesizing Relational Databases from Unstructured Text. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 31987–32012, Suzhou, China. Association for Computational Linguistics.
|
|
| 11 |
Yoshihiko Suhara, Jinfeng Li, Yuliang Li, Dan Zhang, Çağatay Demiralp, Chen Chen, and Wang-Chiew Tan. 2022. Annotating Columns with Pre-trained Language Models. In Proceedings of the 2022 International Conference on Management of Data (SIGMOD '22). Association for Computing Machinery, New York, NY, USA, 1493–1503. https://doi.org/10.1145/3514221.3517906
|
|
| 12 |
Michael Theologitis, Preetam Prabhu Srikar Dammu, Chirag Shah, and Dan Suciu. 2026. ClaimDB: A Fact Verification Benchmark over Large Structured Data. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 34428–34451, San Diego, California, United States. Association for Computational Linguistics.
|
|
| 13 |
Naeemul Hassan, Fatma Arslan, Chengkai Li, and Mark Tremayne. 2017. Toward Automated Fact-Checking: Detecting Check-worthy Factual Claims by ClaimBuster. In Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD '17). Association for Computing Machinery, New York, NY, USA, 1803–1812. https://doi.org/10.1145/3097983.3098131
|
|
| 14 |
Pengcheng Yin, Graham Neubig, Wen-tau Yih, and Sebastian Riedel. 2020. TaBERT: Pretraining for Joint Understanding of Textual and Tabular Data. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8413–8426, Online. Association for Computational Linguistics.
|
|
| 15 |
Tao Yu, Rui Zhang, Kai Yang, Michihiro Yasunaga, Dongxu Wang, Zifan Li, James Ma, Irene Li, Qingning Yao, Shanelle Roman, Zilin Zhang, and Dragomir Radev. 2018. Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 3911–3921, Brussels, Belgium. Association for Computational Linguistics.
|
|
| 16 |
Tianping Zhang, Shaowen Wang, Shuicheng Yan, Li Jian, and Qian Liu. 2023. Generative Table Pre-training Empowers Models for Tabular Prediction. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 14836–14854, Singapore. Association for Computational Linguistics.
|
|