Robuta

https://openreview.net/forum?id=Huw15LqglI Born a Transformer -- Always a Transformer? On the Effect of Pretraining on Architectural Abilities... Transformers have theoretical limitations in modeling certain sequence-to-sequence tasks, yet it remains largely unclear if these limitations play a role in... always onthe effectborntransformer https://www.amazon.science/publications/self-supervised-pretraining-for-large-scale-point-clouds Self-supervised pretraining for large-scale point clouds - Amazon Science Pretraining on large unlabeled datasets has been proven to improve the down stream task performance on many computer vision tasks, such as 2D object detection... self supervisedlarge scalepoint cloudspretrainingamazon https://aclanthology.org/2023.acl-long.656/ From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political... Shangbin Feng, Chan Young Park, Yuhan Liu, Yulia Tsvetkov. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1:... https://arxiv.org/html/2512.06104v1 ARC-AGI Without Pretraining arc agiwithoutpretraining https://huggingface.co/papers/2505.22232 Paper page - Judging Quality Across Languages: A Multilingual Approach to Pretraining Data... Join the discussion on this paper page https://oecd.ai/en/catalogue/metric-use-cases/mtp-advancing-remote-sensing-foundation-model-via-multi-task-pretraining-1 MTP: Advancing Remote Sensing Foundation Model via Multi-Task Pretraining - OECD.AI We propose a novel model-selection method for dynamic real-life networks. Our approach involves training a classifier on a large body of synthetic network... remote sensingfoundation model https://arxiv.org/abs/2101.11363 [2101.11363] KoreALBERT: Pretraining a Lite BERT Model for Korean Language Understanding Abstract page for arXiv paper 2101.11363: KoreALBERT: Pretraining a Lite BERT Model for Korean Language Understanding https://openreview.net/forum?id=HP7Qpui5YE Understanding Self-Supervised Pretraining with Part-Aware Representation Learning | OpenReview In this paper, we are interested in understanding self-supervised pretraining through studying the capability that self-supervised methods learn part-aware... understanding selfrepresentation learningsupervisedpretrainingpart https://aclanthology.org/2025.acl-long.88/ LangSAMP: Language-Script Aware Multilingual Pretraining - ACL Anthology Yihong Liu, Haotian Ye, Chunlan Ma, Mingyang Wang, Hinrich Schuetze. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics... language scriptawaremultilingualpretrainingacl https://openreview.net/forum?id=ssWi0rC3mx When Priors Backfire: On the Vulnerability of Unlearnable Examples to Pretraining | OpenReview Unlearnable Examples (UEs) serve as a data protection strategy that generates imperceptible perturbations to mislead models into learning spurious correlations... backfire on https://openreview.net/forum?id=otHhLO7GZj BACKTRACKING MATHEMATICAL REASONING OF LANGUAGE MODELS TO THE PRETRAINING DATA | OpenReview In this study, we identify subsets of model pretraining data that contribute to the math reasoning ability of language models and evaluate it on several... mathematical reasoninglanguage modelsbacktracking https://openreview.net/forum?id=n0wxja97mN&referrer=%5Bthe%20profile%20of%20Jeonghwan%20Cheon%5D(%2Fprofile%3Fid%3D~Jeonghwan_Cheon1) Pretraining with random noise for uncertainty calibration | OpenReview Uncertainty calibration is crucial for various machine learning applications, yet it remains challenging. Many models exhibit hallucinations - confident yet... random noisepretraininguncertaintycalibrationopenreview https://allenai.org/blog/critical-batch-size Revisiting critical batch size for large-batch OLMo pretraining | Ai2 We introduce a more reliable method to measure the critical batch size (CBS), analyze how CBS changes over training, and use this to train OLMo with fewer grad... batch sizerevisitingcriticallargeolmo https://openreview.net/forum?id=lRvV9rcAbda Generative Pretraining for Black-Box Optimization | OpenReview Many problems in science and engineering involve optimizing an expensive black-box function over a high-dimensional space. For such black-box optimization... black boxgenerativepretrainingoptimizationopenreview https://openreview.net/forum?id=o5z9Le5drua Pretraining Reward-Free Representations for Data-Efficient Reinforcement Learning | OpenReview We show that pre-training encoders with a set of self-supervised learning tasks greatly improves performance in data-efficient RL. for datareinforcement learningpretrainingrewardfree https://openreview.net/forum?id=EDoD3DgivF On Linear Representations and Pretraining Data Frequency in Language Models | OpenReview Pretraining data has a direct impact on the behaviors and quality of language models (LMs), but we only understand the most basic principles of this... linear representationslanguage modelspretraining https://openreview.net/forum?id=gkyosluSbR A Practitioner's Guide to Continual Multimodal Pretraining | OpenReview Multimodal foundation models, despite being extensively pretrained, become outdated over time. Research into continual pretraining mainly explores (1)... guide topractitionercontinualmultimodalpretraining https://openreview.net/forum?id=NCOP0KYb0u&referrer=%5Bthe%20profile%20of%20Byeongguk%20Jeon%5D(%2Fprofile%3Fid%3D~Byeongguk_Jeon1) Latent Action Pretraining From Videos | OpenReview We introduce Latent Action Pretraining for general Action models (LAPA), the first unsupervised method for pretraining Vision-Language-Action (VLA) models... latentactionpretrainingvideosopenreview https://openreview.net/forum?id=5IH0pideQK In-Context Multi-Armed Bandits via Supervised Pretraining | OpenReview Exploring the in-context learning capabilities of large transformer models, this research focuses on decision-making within reinforcement learning (RL)... multi armed banditsin contextviasupervisedpretraining https://arxiv.org/abs/2010.12547v2 [2010.12547v2] Multilingual BERT Post-Pretraining Alignment Abstract page for arXiv paper 2010.12547v2: Multilingual BERT Post-Pretraining Alignment 2010multilingualbertpostpretraining https://aclanthology.org/2022.bionlp-1.9/ BioBART: Pretraining and Evaluation of A Biomedical Generative Language Model - ACL Anthology Hongyi Yuan, Zheng Yuan, Ruyi Gan, Jiaxing Zhang, Yutao Xie, Sheng Yu. Proceedings of the 21st Workshop on Biomedical Language Processing. 2022. of a https://openreview.net/forum?id=Bkl87h09FX Looking for ELMo's friends: Sentence-Level Pretraining Beyond Language Modeling | OpenReview We compare many tasks and task combinations for pretraining sentence-level BiLSTMs for NLP tasks. Language modeling is the best single pretraining task, but... looking for https://openreview.net/forum?id=TRKwzPnXWQ ARVideo: Autoregressive Pretraining for Self-Supervised Video Representation Learning | OpenReview self supervisedrepresentation learningautoregressivepretrainingvideo https://huggingface.co/papers/2504.16511 Paper page - QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining Join the discussion on this paper page data selection for https://arxiv.org/abs/2504.16980 [2504.16980] Safety Pretraining: Toward the Next Generation of Safe AI Abstract page for arXiv paper 2504.16980: Safety Pretraining: Toward the Next Generation of Safe AI the next generationsafetypretrainingtoward https://openreview.net/forum?id=ncjhi4qAPV Position: Considerations for Differentially Private Learning with Large-Scale Public Pretraining |... The performance of differentially private machine learning can be boosted significantly by leveraging the transfer learning capabilities of non-private models... large scalepositionconsiderationsprivate https://deepai.org/publication/3d-denoisers-are-good-2d-teachers-molecular-pretraining-via-denoising-and-cross-modal-distillation 3D Denoisers are Good 2D Teachers: Molecular Pretraining via Denoising and Cross-Modal Distillation... Sep 8, 2023 - 09/08/23 - Pretraining molecular representations from large unlabeled data is essential for molecular property prediction due to the high cos... https://aclanthology.org/2023.ijcnlp-main.8/ On a Benefit of Masked Language Model Pretraining: Robustness to Simplicity Bias - ACL Anthology Ting-Rui Chiang. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of... https://openreview.net/forum?id=ZkFuUac0Hc0 Understanding and Improving Sequence-to-Sequence Pretraining for Neural Machine Translation |... In this paper, we present a substantial step in better understanding the SOTA sequence-to-sequence (Seq2Seq) pretraining for neural machine translation~(NMT).... understandingimprovingsequencepretrainingneural https://arxiv.org/abs/2404.06214v2 [2404.06214v2] [Call for Papers] The 2nd BabyLM Challenge: Sample-efficient pretraining on a... Abstract page for arXiv paper 2404.06214v2: [Call for Papers] The 2nd BabyLM Challenge: Sample-efficient pretraining on a developmentally plausible corpus https://indiandefencepretraining.com/ Indian defence pretraining indian defencepretraining https://openreview.net/forum?id=8yRnzgIADf&referrer=%5Bthe%20profile%20of%20Guang%20Liu%5D(%2Fprofile%3Fid%3D~Guang_Liu2) UniTabE: A Universal Pretraining Protocol for Tabular Foundation Model in Data Science | OpenReview https://openreview.net/forum?id=EUfUgS17T9 LinkBERT: Pretraining Language Models with Document Links | OpenReview We propose LinkBERT, a new language model pretraining method that incorporates document link information (e.g. hyperlinks, citation links), and show that it... language modelsdocument linkspretrainingopenreview https://deepai.org/publication/guiding-pretraining-in-reinforcement-learning-with-large-language-models Guiding Pretraining in Reinforcement Learning with Large Language Models | DeepAI Feb 13, 2023 - 02/13/23 - Reinforcement learning algorithms typically struggle in the absence of a dense, well-shaped reward function. Intrinsically motivat... large language modelsreinforcement learningguidingpretrainingdeepai https://aclanthology.org/2022.findings-naacl.32/ DOCmT5: Document-Level Pretraining of Multilingual Language Models - ACL Anthology Chia-Hsuan Lee, Aditya Siddhant, Viresh Ratnakar, Melvin Johnson. Findings of the Association for Computational Linguistics: NAACL 2022. 2022. multilingual languagedocumentlevelpretrainingmodels https://deepai.org/publication/paradise-exploiting-parallel-data-for-multilingual-sequence-to-sequence-pretraining PARADISE: Exploiting Parallel Data for Multilingual Sequence-to-Sequence Pretraining | DeepAI Aug 4, 2021 - 08/04/21 - Despite the success of multilingual sequence-to-sequence pretraining, most existing approaches rely on monolingual corpora, and do... data forparadiseexploitingparallelmultilingual https://arxiv.org/abs/2111.04130?ref=dataphoenix.info [2111.04130] NLP From Scratch Without Large-Scale Pretraining: A Simple and Efficient Framework Abstract page for arXiv paper 2111.04130: NLP From Scratch Without Large-Scale Pretraining: A Simple and Efficient Framework https://deepai.org/publication/convergence-rates-for-pretraining-and-dropout-guiding-learning-parameters-using-network-structure Convergence rates for pretraining and dropout: Guiding learning parameters using network structure... Jun 10, 2015 - 06/10/15 - Unsupervised pretraining and dropout have been well studied, especially with respect to regularization and output consistency. How... https://openreview.net/forum?id=PQpvhUrA1C Autoregressive Pretraining with Mamba in Vision | OpenReview The vision community has started to build with the recently developed state space model, Mamba, as the new backbone for a range of tasks. This paper shows that... in visionautoregressivepretrainingmambaopenreview https://openreview.net/forum?id=N7DSUbnzYo&referrer=%5BAuthor%20Console%5D(%2Fgroup%3Fid%3DTMLR%2FAuthors%23your-submissions) Language-Pretraining-Induced Bias: A Strong Foundation for General Vision Tasks | OpenReview The ratio of "outlier" parameters in language pre-training models and vision pre-training models differs significantly, making cross-modality (language and... a strong https://castbox.fm/episode/Pretraining-AI-id7055509-id907329080 Pretraining AI pretraining https://research.google/pubs/end-to-end-generative-pretraining-for-multimodal-video-captioning/ End-to-end Generative Pretraining for Multimodal Video Captioning endgenerativepretrainingmultimodalvideo https://openreview.net/forum?id=XrsOu4KgDE Attributing Culture-Conditioned Generations to Pretraining Corpora | OpenReview In open-ended generative tasks like narrative writing or dialogue, large language models often exhibit cultural biases, showing limited knowledge and... attributingcultureconditionedgenerationspretraining https://www.amazon.science/publications/graphire-novel-intent-discovery-with-pretraining-on-prior-knowledge-using-contrastive-learning Graphire: Novel Intent Discovery with Pretraining on Prior Knowledge using Contrastive Learning -... In this paper, we introduce Graphire, an intent discovery system leveraging pretraining on predefined intents to automatically discover novel intents for... https://arxiv.org/html/2411.12580v1 Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models procedural knowledgelarge languagepretrainingdrivesreasoning https://warwick.ac.uk/fac/cross_fac/eduport/edufund/projects/yang/projects/safety-pretraining-toward-the-next-generation-of-safe-ai/ safety-pretraining-toward-the-next-generation-of-safe-ai Safety Pretraining: Toward the Next Generation of Safe AI - Research project on AI in education the next generationsafetypretrainingtoward https://www.amazon.science/publications/ceres-pretraining-of-graph-conditioned-transformer-for-semi-structured-session-data CERES: Pretraining of graph-conditioned transformer for semi-structured session data - Amazon... User sessions empower many search and recommendation tasks on a daily basis. Such session data are semi-structured, which encode heterogeneous relations... https://openreview.net/forum?id=1hQKHHUsMx&referrer=%5Bthe%20profile%20of%20Max%20Bartolo%5D(%2Fprofile%3Fid%3D~Max_Bartolo1) Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models | OpenReview The capabilities and limitations of Large Language Models (LLMs) have been sketched out in great detail in recent years, providing an intriguing yet... large language modelsprocedural knowledgepretrainingdrivesreasoning https://cohere.com/research/papers/investigating-continual-pretraining-in-large-language-models-insights-and-implications-2024-02-27 Investigating Continual Pretraining in Large Language Models: Insights and Implications This paper studies the evolving domain of Continual Learning (CL) in large language models (LLMs), with a focus on developing strategies for efficient and large language modelsinvestigatingcontinualpretraininginsights https://www.amazon.science/publications/how-much-pretraining-data-do-language-models-need-to-learn-syntax How much pretraining data do language models need to learn syntax? - Amazon Science Transformers-based pretrained language models achieve outstanding results in many well-known NLU benchmarks. However, while pretraining methods are very... https://huggingface.co/papers/2401.03003 Paper page - AST-T5: Structure-Aware Pretraining for Code Generation and Understanding Join the discussion on this paper page https://arxiv.org/abs/2505.14683 [2505.14683] Emerging Properties in Unified Multimodal Pretraining Abstract page for arXiv paper 2505.14683: Emerging Properties in Unified Multimodal Pretraining emerging propertiesunifiedmultimodalpretraining https://arxiv.org/abs/2002.01685 [2002.01685] Parsing as Pretraining Abstract page for arXiv paper 2002.01685: Parsing as Pretraining 2002parsingpretraining https://aclanthology.org/2025.babylm-main.33/ You are an LLM teaching a smaller model everything you know: Multi-task pretraining of language... Wiktor Kamzela, Mateusz Lango, Ondrej Dusek. Proceedings of the First BabyLM Workshop. 2025. https://www.mdpi.com/1424-8220/22/17/6504 Visual Pretraining via Contrastive Predictive Model for Pixel-Based Reinforcement Learning In an attempt to overcome the limitations of reward-driven representation learning in vision-based reinforcement learning (RL), an unsupervised learning... predictive modelvisualpretrainingviacontrastive https://www.themuse.com/jobs/servicenow/staff-applied-research-scientist-pretrainingfinetuning Staff Applied Research Scientist - Pretraining/Fine-tuning at ServiceNow | The Muse | The Muse Find our Staff Applied Research Scientist - Pretraining/Fine-tuning job description for ServiceNow located in Santa Clara, CA, as well as other career... applied researchfine tuningstaffscientistpretraining https://deepai.org/publication/don-t-stop-pretraining-adapt-language-models-to-domains-and-tasks Don't Stop Pretraining: Adapt Language Models to Domains and Tasks | DeepAI Apr 23, 2020 - 04/23/20 - Language models pretrained on text from a wide variety of sources form the foundation of today's NLP. In light of the success of t... don t stoplanguage models https://github.com/facebookresearch/ssvp_slt GitHub - facebookresearch/ssvp_slt: Self-supervised video pretraining for sign language... Self-supervised video pretraining for sign language translation. - facebookresearch/ssvp_slt self supervisedgithubfacebookresearchssvpslt https://research.google/pubs/analyzing-similarity-metrics-for-data-selection-for-language-model-pretraining/ Analyzing Similarity Metrics for Data Selection for Language Model Pretraining for datalanguage modelanalyzingsimilaritymetrics https://openreview.net/forum?id=N7DSUbnzYo Language-Pretraining-Induced Bias: A Strong Foundation for General Vision Tasks | OpenReview The ratio of "outlier" parameters in language pre-training models and vision pre-training models differs significantly, making cross-modality (language and... a strong https://aclanthology.org/2020.findings-emnlp.414/ Byte Pair Encoding is Suboptimal for Language Model Pretraining - ACL Anthology Kaj Bostrom, Greg Durrett. Findings of the Association for Computational Linguistics: EMNLP 2020. 2020. byte pair encodingfor languagesuboptimal https://arxiv.org/abs/2512.06104v1 [2512.06104v1] ARC-AGI Without Pretraining Abstract page for arXiv paper 2512.06104v1: ARC-AGI Without Pretraining arc agi2512withoutpretraining https://velog.io/@sangwu99/Pix2Struct-Screenshot-Parsing-as-Pretraining-for-Visual-Language-Understanding-ICML-2023 Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding (ICML 2023) visual languagescreenshotparsingpretraining https://arxiv.org/html/2604.05215v1 Hierarchical Mesh Transformers with Topology-Guided Pretraining for Morphometric Analysis of Brain... https://arxiv.org/html/2410.05612v2 A Bayesian Model Selection Criterion for Selecting Pretraining Checkpoints bayesian model selectioncriterionselectingpretrainingcheckpoints https://openreview.net/forum?id=sqBIm0Irju7 Surprisingly Simple Semi-Supervised Domain Adaptation with Pretraining and Consistency | OpenReview A strong method for semi-supervised domain adaptation that unlike most recent state of the art approaches does not rely on explicit domain alignment domain adaptationsurprisinglysimplesemisupervised https://arxiv.org/abs/2506.22049 [2506.22049] GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation... Abstract page for arXiv paper 2506.22049: GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling https://openreview.net/forum?id=7qfkImn0dL ExPT: Synthetic Pretraining for Few-Shot Experimental Design | OpenReview Experimental design is a fundamental problem in many science and engineering fields. In this problem, sample efficiency is crucial due to the time, money, and... few shotexperimental designsyntheticpretrainingopenreview https://oecd.ai/en/catalogue/metric-use-cases/mtp-advancing-remote-sensing-foundation-model-via-multi-task-pretraining-3 MTP: Advancing Remote Sensing Foundation Model via Multi-Task Pretraining - OECD.AI We propose a novel model-selection method for dynamic real-life networks. Our approach involves training a classifier on a large body of synthetic network... remote sensingfoundation model https://openreview.net/forum?id=hLGJ1qZPdu On the Generalization Ability of Next-Token-Prediction Pretraining | OpenReview Large language models (LLMs) have demonstrated remarkable potential in handling natural language processing (NLP) tasks and beyond. LLMs usually can be... on thegeneralizationabilitynexttoken https://allenai.org/blog/dolma-3-trillion-tokens-open-llm-corpus-9a0ff4b8da64 Ai2 Dolma: 3 trillion token open corpus for language model pretraining | Ai2 We introduce Dolma, an open dataset from web content, academic publications, code, books, and encyclopedic materials. for languageai2dolma3trillion https://arxiv.org/abs/2310.03291v1 [2310.03291v1] SimVLG: Simple and Efficient Pretraining of Visual Language Generative Models Abstract page for arXiv paper 2310.03291v1: SimVLG: Simple and Efficient Pretraining of Visual Language Generative Models https://allenai.org/blog/datadecide DataDecide: How to predict best pretraining data with small experiments | Ai2 Explore the secrets of how language model developers make decisions with DataDecide. how topredictbest https://openreview.net/forum?id=ZeFMtRBy4Z REVE: A Foundation Model for EEG - Adapting to Any Setup with Large-Scale Pretraining on 25,000... Foundation models have transformed AI by reducing reliance on task-specific data through large-scale pretraining. While successful in language and vision,... https://openreview.net/forum?id=XlAbMZu4Bo&referrer=%5Bthe%20profile%20of%20Wenhan%20Xiong%5D(%2Fprofile%3Fid%3D~Wenhan_Xiong1) Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context Length | OpenReview The quadratic complexity and weak length extrapolation of Transformers limits their ability to scale to long sequences, and while sub-quadratic solutions like... context lengthmegalodonefficientllmpretraining https://deepai.org/publication/chinese-clip-contrastive-vision-language-pretraining-in-chinese Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese | DeepAI Nov 2, 2022 - 11/02/22 - The tremendous success of CLIP (Radford et al., 2021) has promoted the research and application of contrastive learning for vision... vision languagechineseclipcontrastivepretraining https://www.sydney.edu.au/scholarships/b/vision-language-pretraining-model.html Postgraduate Research Scholarship in Vision Language Pretraining (VLP) model - Scholarships postgraduate researchin visionscholarshiplanguagepretraining https://openreview.net/forum?id=QOcukxDfTq DynaMo: In-Domain Dynamics Pretraining for Visuo-Motor Control | OpenReview Imitation learning has proven to be a powerful tool for training complex visuo-motor policies. However, current methods often require hundreds to thousands of... in domainvisuo motordynamodynamicspretraining https://openreview.net/forum?id=TYdzj1EvBP How Do Large Language Models Acquire Factual Knowledge During Pretraining? | OpenReview Despite the recent observation that large language models (LLMs) can store substantial factual knowledge, there is a limited understanding of the mechanisms of... large language modelshow dofactual knowledge https://deepai.org/publication/protein-representation-learning-by-geometric-structure-pretraining Protein Representation Learning by Geometric Structure Pretraining | DeepAI Mar 11, 2022 - 03/11/22 - Learning effective protein representations is critical in a variety of tasks in biology such as predicting protein function or str... representation learninggeometric structureproteinpretrainingdeepai https://openreview.net/forum?id=YhWFvZcahs&referrer=%5Bthe%20profile%20of%20Fabian%20Caba%20Heilbron%5D(%2Fprofile%3Fid%3D~Fabian_Caba_Heilbron3) Scaling Up Video Summarization Pretraining with Large Language Models | OpenReview Long-form video content constitutes a significant portion of internet traffic, making automated video summarization an essential research problem. However,... large language modelsscaling upvideo summarizationpretrainingopenreview https://openreview.net/forum?id=sV4Y6i1lvE&referrer=%5Bthe%20profile%20of%20Jianwei%20Yang%5D(%2Fprofile%3Fid%3D~Jianwei_Yang1) Latent Action Pretraining From Videos | OpenReview We introduce Latent Action Pretraining for general Action models (LAPA), the first unsupervised method for pretraining Vision-Language-Action (VLA) models... latentactionpretrainingvideosopenreview https://aclanthology.org/2020.aacl-main.24/ UnihanLM: Coarse-to-Fine Chinese-Japanese Language Model Pretraining with the Unihan Database - ACL... Canwen Xu, Tao Ge, Chenliang Li, Furu Wei. Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and... https://www.amazon.science/publications/multitask-pretraining-with-structured-knowledge-for-text-to-sql-generation Multitask pretraining with structured knowledge for text-to-SQL generation - Amazon Science Many machine learning-based low-code or no-code applications involve generating code that interacts with structured knowledge. For example, one of the most... text to sql https://deepai.org/publication/structure-grounded-pretraining-for-text-to-sql Structure-Grounded Pretraining for Text-to-SQL | DeepAI Oct 24, 2020 - 10/24/20 - Learning to capture text-table alignment is essential for table related tasks like text-to-SQL. The model needs to correctly recog... text to sqlstructuregroundedpretrainingdeepai https://www.amazon.science/blog/building-geospatial-foundation-models-via-continual-pretraining Building geospatial foundation models via continual pretraining - Amazon Science Mar 26, 2024 - New approach enables sustainable machine learning for remote-sensing applications. foundation modelsbuildinggeospatialviacontinual https://openreview.net/forum?id=CjTHVo1dvR Molecular Geometry Pretraining with SE(3)-Invariant Denoising Distance Matching | OpenReview We propose GeoSSL, a self-supervised learning method using the denoising distance matching for molecular goemetry pretraining. molecular geometrypretrainingse https://openreview.net/forum?id=ukgVYJBkvO Learning the Neighborhood: Contrast-Free Self-Supervised Molecular Graph Pretraining | OpenReview High-quality molecular representations are essential for property prediction and molecular design, yet large labeled datasets remain scarce. Self-supervised... the neighborhoodself supervisedmolecular graphlearningcontrast https://aclanthology.org/2023.unlp-1.4/ GPT-2 Metadata Pretraining Towards Instruction Finetuning for Ukrainian - ACL Anthology Volodymyr Kyrylov, Dmytro Chaplynskyi. Proceedings of the Second Ukrainian Natural Language Processing Workshop (UNLP). 2023. gpt 2metadatapretrainingtowards https://aclanthology.org/2024.knowllm-1.9/ Knowledge Acquisition through Continued Pretraining is Difficult: A Case Study on r/AskHistorians -... Jan Hoffbauer, Sylwester Sawicki, Marc Ulrich, Tolga Buz, Konstantin Dobler, Moritz Schneider, Gerard De Melo. Proceedings of the 1st Workshop on Towards... a case study https://openreview.net/forum?id=1Iuw1jcIrf MathCoder2: Better Math Reasoning from Continued Pretraining on Model-translated Mathematical Code... Code has been shown to be effective in enhancing the mathematical reasoning abilities of large language models due to its precision and accuracy. Previous... https://arxiv.org/abs/2505.02009 [2505.02009] Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale... Abstract page for arXiv paper 2505.02009: Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs https://deepai.org/publication/why-is-public-pretraining-necessary-for-private-model-training Why Is Public Pretraining Necessary for Private Model Training? | DeepAI Feb 19, 2023 - 02/19/23 - In the privacy-utility tradeoff of a model trained on benchmark language and vision tasks, remarkable improvements have been widel... for privatemodel trainingpublicpretrainingnecessary https://openreview.net/forum?id=B8BXHrshMi FlexRibbon: Joint Sequence and Structure Pretraining for Protein Modeling | OpenReview Protein foundation models have advanced rapidly, with most approaches falling into two dominant paradigms. Sequence-based language models (e.g., ESM-2) capture... jointsequencestructurepretrainingprotein https://openreview.net/forum?id=VYOe2eBQeh Latent Action Pretraining from Videos | OpenReview We introduce Latent Action Pretraining for general Action models (LAPA), the first unsupervised method for pretraining Vision-Language-Action (VLA) models... latentactionpretrainingvideosopenreview https://openreview.net/forum?id=1mjsP8RYAw Unsupervised Pretraining for Fact Verification by Language Model Distillation | OpenReview Fact verification aims to verify a claim using evidence from a trustworthy knowledge base. To address this challenge, algorithms must produce features for... by languagemodel distillationunsupervisedpretrainingfact https://aclanthology.org/2025.babylm-main.37/ Pretraining Language Models with LoRA and Artificial Languages - ACL Anthology Nalin Kumar, Mateusz Lango, Ondrej Dusek. Proceedings of the First BabyLM Workshop. 2025. language modelsartificial languagespretrainingloraacl https://deepai.org/publication/on-the-effective-use-of-pretraining-for-natural-language-inference On the Effective Use of Pretraining for Natural Language Inference | DeepAI Oct 5, 2017 - 10/05/17 - Neural networks have excelled at many NLP tasks, but there remain open questions about the performance of pretrained distributed w... natural language inferenceon theeffective use https://huggingface.co/papers/2406.02214 Paper page - SLTrain: a sparse plus low-rank approach for parameter and memory efficient pretraining Join the discussion on this paper page https://openreview.net/forum?id=zBBmV-i84Go Addressing Resource Scarcity across Sign Languages with Multilingual Pretraining and... We release the largest available pretraining dataset for sign language across multiple languages and show how multilingual fine-tuning using a unified... resource scarcitysign languagesaddressingacrossmultilingual