https://openreview.net/forum?id=Huw15LqglI
Born a Transformer -- Always a Transformer? On the Effect of Pretraining on Architectural Abilities...
Transformers have theoretical limitations in modeling certain sequence-to-sequence tasks, yet it remains largely unclear if these limitations play a role in...
always onthe effectborntransformer
https://www.amazon.science/publications/self-supervised-pretraining-for-large-scale-point-clouds
Self-supervised pretraining for large-scale point clouds - Amazon Science
Pretraining on large unlabeled datasets has been proven to improve the down stream task performance on many computer vision tasks, such as 2D object detection...
self supervisedlarge scalepoint cloudspretrainingamazon
https://aclanthology.org/2023.acl-long.656/
From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political...
Shangbin Feng, Chan Young Park, Yuhan Liu, Yulia Tsvetkov. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1:...
https://arxiv.org/html/2512.06104v1
ARC-AGI Without Pretraining
arc agiwithoutpretraining
https://huggingface.co/papers/2505.22232
Paper page - Judging Quality Across Languages: A Multilingual Approach to Pretraining Data...
Join the discussion on this paper page
https://oecd.ai/en/catalogue/metric-use-cases/mtp-advancing-remote-sensing-foundation-model-via-multi-task-pretraining-1
MTP: Advancing Remote Sensing Foundation Model via Multi-Task Pretraining - OECD.AI
We propose a novel model-selection method for dynamic real-life networks. Our approach involves training a classifier on a large body of synthetic network...
remote sensingfoundation model
https://arxiv.org/abs/2101.11363
[2101.11363] KoreALBERT: Pretraining a Lite BERT Model for Korean Language Understanding
Abstract page for arXiv paper 2101.11363: KoreALBERT: Pretraining a Lite BERT Model for Korean Language Understanding
https://openreview.net/forum?id=HP7Qpui5YE
Understanding Self-Supervised Pretraining with Part-Aware Representation Learning | OpenReview
In this paper, we are interested in understanding self-supervised pretraining through studying the capability that self-supervised methods learn part-aware...
understanding selfrepresentation learningsupervisedpretrainingpart
https://aclanthology.org/2025.acl-long.88/
LangSAMP: Language-Script Aware Multilingual Pretraining - ACL Anthology
Yihong Liu, Haotian Ye, Chunlan Ma, Mingyang Wang, Hinrich Schuetze. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics...
language scriptawaremultilingualpretrainingacl
https://openreview.net/forum?id=ssWi0rC3mx
When Priors Backfire: On the Vulnerability of Unlearnable Examples to Pretraining | OpenReview
Unlearnable Examples (UEs) serve as a data protection strategy that generates imperceptible perturbations to mislead models into learning spurious correlations...
backfire on
https://openreview.net/forum?id=otHhLO7GZj
BACKTRACKING MATHEMATICAL REASONING OF LANGUAGE MODELS TO THE PRETRAINING DATA | OpenReview
In this study, we identify subsets of model pretraining data that contribute to the math reasoning ability of language models and evaluate it on several...
mathematical reasoninglanguage modelsbacktracking
https://openreview.net/forum?id=n0wxja97mN&referrer=%5Bthe%20profile%20of%20Jeonghwan%20Cheon%5D(%2Fprofile%3Fid%3D~Jeonghwan_Cheon1)
Pretraining with random noise for uncertainty calibration | OpenReview
Uncertainty calibration is crucial for various machine learning applications, yet it remains challenging. Many models exhibit hallucinations - confident yet...
random noisepretraininguncertaintycalibrationopenreview
https://allenai.org/blog/critical-batch-size
Revisiting critical batch size for large-batch OLMo pretraining | Ai2
We introduce a more reliable method to measure the critical batch size (CBS), analyze how CBS changes over training, and use this to train OLMo with fewer grad...
batch sizerevisitingcriticallargeolmo
https://openreview.net/forum?id=lRvV9rcAbda
Generative Pretraining for Black-Box Optimization | OpenReview
Many problems in science and engineering involve optimizing an expensive black-box function over a high-dimensional space. For such black-box optimization...
black boxgenerativepretrainingoptimizationopenreview
https://openreview.net/forum?id=o5z9Le5drua
Pretraining Reward-Free Representations for Data-Efficient Reinforcement Learning | OpenReview
We show that pre-training encoders with a set of self-supervised learning tasks greatly improves performance in data-efficient RL.
for datareinforcement learningpretrainingrewardfree
https://openreview.net/forum?id=EDoD3DgivF
On Linear Representations and Pretraining Data Frequency in Language Models | OpenReview
Pretraining data has a direct impact on the behaviors and quality of language models (LMs), but we only understand the most basic principles of this...
linear representationslanguage modelspretraining
https://openreview.net/forum?id=gkyosluSbR
A Practitioner's Guide to Continual Multimodal Pretraining | OpenReview
Multimodal foundation models, despite being extensively pretrained, become outdated over time. Research into continual pretraining mainly explores (1)...
guide topractitionercontinualmultimodalpretraining
https://openreview.net/forum?id=NCOP0KYb0u&referrer=%5Bthe%20profile%20of%20Byeongguk%20Jeon%5D(%2Fprofile%3Fid%3D~Byeongguk_Jeon1)
Latent Action Pretraining From Videos | OpenReview
We introduce Latent Action Pretraining for general Action models (LAPA), the first unsupervised method for pretraining Vision-Language-Action (VLA) models...
latentactionpretrainingvideosopenreview
https://openreview.net/forum?id=5IH0pideQK
In-Context Multi-Armed Bandits via Supervised Pretraining | OpenReview
Exploring the in-context learning capabilities of large transformer models, this research focuses on decision-making within reinforcement learning (RL)...
multi armed banditsin contextviasupervisedpretraining
https://arxiv.org/abs/2010.12547v2
[2010.12547v2] Multilingual BERT Post-Pretraining Alignment
Abstract page for arXiv paper 2010.12547v2: Multilingual BERT Post-Pretraining Alignment
2010multilingualbertpostpretraining
https://aclanthology.org/2022.bionlp-1.9/
BioBART: Pretraining and Evaluation of A Biomedical Generative Language Model - ACL Anthology
Hongyi Yuan, Zheng Yuan, Ruyi Gan, Jiaxing Zhang, Yutao Xie, Sheng Yu. Proceedings of the 21st Workshop on Biomedical Language Processing. 2022.
of a
https://openreview.net/forum?id=Bkl87h09FX
Looking for ELMo's friends: Sentence-Level Pretraining Beyond Language Modeling | OpenReview
We compare many tasks and task combinations for pretraining sentence-level BiLSTMs for NLP tasks. Language modeling is the best single pretraining task, but...
looking for
https://openreview.net/forum?id=TRKwzPnXWQ
ARVideo: Autoregressive Pretraining for Self-Supervised Video Representation Learning | OpenReview
self supervisedrepresentation learningautoregressivepretrainingvideo
https://huggingface.co/papers/2504.16511
Paper page - QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining
Join the discussion on this paper page
data selection for
https://arxiv.org/abs/2504.16980
[2504.16980] Safety Pretraining: Toward the Next Generation of Safe AI
Abstract page for arXiv paper 2504.16980: Safety Pretraining: Toward the Next Generation of Safe AI
the next generationsafetypretrainingtoward
https://openreview.net/forum?id=ncjhi4qAPV
Position: Considerations for Differentially Private Learning with Large-Scale Public Pretraining |...
The performance of differentially private machine learning can be boosted significantly by leveraging the transfer learning capabilities of non-private models...
large scalepositionconsiderationsprivate
https://deepai.org/publication/3d-denoisers-are-good-2d-teachers-molecular-pretraining-via-denoising-and-cross-modal-distillation
3D Denoisers are Good 2D Teachers: Molecular Pretraining via Denoising and Cross-Modal Distillation...
Sep 8, 2023 - 09/08/23 - Pretraining molecular representations from large unlabeled data is essential for molecular property prediction due to the high cos...
https://aclanthology.org/2023.ijcnlp-main.8/
On a Benefit of Masked Language Model Pretraining: Robustness to Simplicity Bias - ACL Anthology
Ting-Rui Chiang. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of...
https://openreview.net/forum?id=ZkFuUac0Hc0
Understanding and Improving Sequence-to-Sequence Pretraining for Neural Machine Translation |...
In this paper, we present a substantial step in better understanding the SOTA sequence-to-sequence (Seq2Seq) pretraining for neural machine translation~(NMT)....
understandingimprovingsequencepretrainingneural
https://arxiv.org/abs/2404.06214v2
[2404.06214v2] [Call for Papers] The 2nd BabyLM Challenge: Sample-efficient pretraining on a...
Abstract page for arXiv paper 2404.06214v2: [Call for Papers] The 2nd BabyLM Challenge: Sample-efficient pretraining on a developmentally plausible corpus
https://indiandefencepretraining.com/
Indian defence pretraining
indian defencepretraining
https://openreview.net/forum?id=8yRnzgIADf&referrer=%5Bthe%20profile%20of%20Guang%20Liu%5D(%2Fprofile%3Fid%3D~Guang_Liu2)
UniTabE: A Universal Pretraining Protocol for Tabular Foundation Model in Data Science | OpenReview
https://openreview.net/forum?id=EUfUgS17T9
LinkBERT: Pretraining Language Models with Document Links | OpenReview
We propose LinkBERT, a new language model pretraining method that incorporates document link information (e.g. hyperlinks, citation links), and show that it...
language modelsdocument linkspretrainingopenreview
https://deepai.org/publication/guiding-pretraining-in-reinforcement-learning-with-large-language-models
Guiding Pretraining in Reinforcement Learning with Large Language Models | DeepAI
Feb 13, 2023 - 02/13/23 - Reinforcement learning algorithms typically struggle in the absence of a dense, well-shaped reward function. Intrinsically motivat...
large language modelsreinforcement learningguidingpretrainingdeepai
https://aclanthology.org/2022.findings-naacl.32/
DOCmT5: Document-Level Pretraining of Multilingual Language Models - ACL Anthology
Chia-Hsuan Lee, Aditya Siddhant, Viresh Ratnakar, Melvin Johnson. Findings of the Association for Computational Linguistics: NAACL 2022. 2022.
multilingual languagedocumentlevelpretrainingmodels
https://deepai.org/publication/paradise-exploiting-parallel-data-for-multilingual-sequence-to-sequence-pretraining
PARADISE: Exploiting Parallel Data for Multilingual Sequence-to-Sequence Pretraining | DeepAI
Aug 4, 2021 - 08/04/21 - Despite the success of multilingual sequence-to-sequence pretraining, most existing approaches rely on monolingual corpora, and do...
data forparadiseexploitingparallelmultilingual
https://arxiv.org/abs/2111.04130?ref=dataphoenix.info
[2111.04130] NLP From Scratch Without Large-Scale Pretraining: A Simple and Efficient Framework
Abstract page for arXiv paper 2111.04130: NLP From Scratch Without Large-Scale Pretraining: A Simple and Efficient Framework
https://deepai.org/publication/convergence-rates-for-pretraining-and-dropout-guiding-learning-parameters-using-network-structure
Convergence rates for pretraining and dropout: Guiding learning parameters using network structure...
Jun 10, 2015 - 06/10/15 - Unsupervised pretraining and dropout have been well studied, especially with respect to regularization and output consistency. How...
https://openreview.net/forum?id=PQpvhUrA1C
Autoregressive Pretraining with Mamba in Vision | OpenReview
The vision community has started to build with the recently developed state space model, Mamba, as the new backbone for a range of tasks. This paper shows that...
in visionautoregressivepretrainingmambaopenreview
https://openreview.net/forum?id=N7DSUbnzYo&referrer=%5BAuthor%20Console%5D(%2Fgroup%3Fid%3DTMLR%2FAuthors%23your-submissions)
Language-Pretraining-Induced Bias: A Strong Foundation for General Vision Tasks | OpenReview
The ratio of "outlier" parameters in language pre-training models and vision pre-training models differs significantly, making cross-modality (language and...
a strong
https://castbox.fm/episode/Pretraining-AI-id7055509-id907329080
Pretraining AI
pretraining
https://research.google/pubs/end-to-end-generative-pretraining-for-multimodal-video-captioning/
End-to-end Generative Pretraining for Multimodal Video Captioning
endgenerativepretrainingmultimodalvideo
https://openreview.net/forum?id=XrsOu4KgDE
Attributing Culture-Conditioned Generations to Pretraining Corpora | OpenReview
In open-ended generative tasks like narrative writing or dialogue, large language models often exhibit cultural biases, showing limited knowledge and...
attributingcultureconditionedgenerationspretraining
https://www.amazon.science/publications/graphire-novel-intent-discovery-with-pretraining-on-prior-knowledge-using-contrastive-learning
Graphire: Novel Intent Discovery with Pretraining on Prior Knowledge using Contrastive Learning -...
In this paper, we introduce Graphire, an intent discovery system leveraging pretraining on predefined intents to automatically discover novel intents for...
https://arxiv.org/html/2411.12580v1
Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models
procedural knowledgelarge languagepretrainingdrivesreasoning
https://warwick.ac.uk/fac/cross_fac/eduport/edufund/projects/yang/projects/safety-pretraining-toward-the-next-generation-of-safe-ai/
safety-pretraining-toward-the-next-generation-of-safe-ai
Safety Pretraining: Toward the Next Generation of Safe AI - Research project on AI in education
the next generationsafetypretrainingtoward
https://www.amazon.science/publications/ceres-pretraining-of-graph-conditioned-transformer-for-semi-structured-session-data
CERES: Pretraining of graph-conditioned transformer for semi-structured session data - Amazon...
User sessions empower many search and recommendation tasks on a daily basis. Such session data are semi-structured, which encode heterogeneous relations...
https://openreview.net/forum?id=1hQKHHUsMx&referrer=%5Bthe%20profile%20of%20Max%20Bartolo%5D(%2Fprofile%3Fid%3D~Max_Bartolo1)
Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models | OpenReview
The capabilities and limitations of Large Language Models (LLMs) have been sketched out in great detail in recent years, providing an intriguing yet...
large language modelsprocedural knowledgepretrainingdrivesreasoning
https://cohere.com/research/papers/investigating-continual-pretraining-in-large-language-models-insights-and-implications-2024-02-27
Investigating Continual Pretraining in Large Language Models: Insights and Implications
This paper studies the evolving domain of Continual Learning (CL) in large language models (LLMs), with a focus on developing strategies for efficient and
large language modelsinvestigatingcontinualpretraininginsights
https://www.amazon.science/publications/how-much-pretraining-data-do-language-models-need-to-learn-syntax
How much pretraining data do language models need to learn syntax? - Amazon Science
Transformers-based pretrained language models achieve outstanding results in many well-known NLU benchmarks. However, while pretraining methods are very...
https://huggingface.co/papers/2401.03003
Paper page - AST-T5: Structure-Aware Pretraining for Code Generation and Understanding
Join the discussion on this paper page
https://arxiv.org/abs/2505.14683
[2505.14683] Emerging Properties in Unified Multimodal Pretraining
Abstract page for arXiv paper 2505.14683: Emerging Properties in Unified Multimodal Pretraining
emerging propertiesunifiedmultimodalpretraining
https://arxiv.org/abs/2002.01685
[2002.01685] Parsing as Pretraining
Abstract page for arXiv paper 2002.01685: Parsing as Pretraining
2002parsingpretraining
https://aclanthology.org/2025.babylm-main.33/
You are an LLM teaching a smaller model everything you know: Multi-task pretraining of language...
Wiktor Kamzela, Mateusz Lango, Ondrej Dusek. Proceedings of the First BabyLM Workshop. 2025.
https://www.mdpi.com/1424-8220/22/17/6504
Visual Pretraining via Contrastive Predictive Model for Pixel-Based Reinforcement Learning
In an attempt to overcome the limitations of reward-driven representation learning in vision-based reinforcement learning (RL), an unsupervised learning...
predictive modelvisualpretrainingviacontrastive
https://www.themuse.com/jobs/servicenow/staff-applied-research-scientist-pretrainingfinetuning
Staff Applied Research Scientist - Pretraining/Fine-tuning at ServiceNow | The Muse | The Muse
Find our Staff Applied Research Scientist - Pretraining/Fine-tuning job description for ServiceNow located in Santa Clara, CA, as well as other career...
applied researchfine tuningstaffscientistpretraining
https://deepai.org/publication/don-t-stop-pretraining-adapt-language-models-to-domains-and-tasks
Don't Stop Pretraining: Adapt Language Models to Domains and Tasks | DeepAI
Apr 23, 2020 - 04/23/20 - Language models pretrained on text from a wide variety of sources form the foundation of today's NLP. In light of the success of t...
don t stoplanguage models
https://github.com/facebookresearch/ssvp_slt
GitHub - facebookresearch/ssvp_slt: Self-supervised video pretraining for sign language...
Self-supervised video pretraining for sign language translation. - facebookresearch/ssvp_slt
self supervisedgithubfacebookresearchssvpslt
https://research.google/pubs/analyzing-similarity-metrics-for-data-selection-for-language-model-pretraining/
Analyzing Similarity Metrics for Data Selection for Language Model Pretraining
for datalanguage modelanalyzingsimilaritymetrics
https://openreview.net/forum?id=N7DSUbnzYo
Language-Pretraining-Induced Bias: A Strong Foundation for General Vision Tasks | OpenReview
The ratio of "outlier" parameters in language pre-training models and vision pre-training models differs significantly, making cross-modality (language and...
a strong
https://aclanthology.org/2020.findings-emnlp.414/
Byte Pair Encoding is Suboptimal for Language Model Pretraining - ACL Anthology
Kaj Bostrom, Greg Durrett. Findings of the Association for Computational Linguistics: EMNLP 2020. 2020.
byte pair encodingfor languagesuboptimal
https://arxiv.org/abs/2512.06104v1
[2512.06104v1] ARC-AGI Without Pretraining
Abstract page for arXiv paper 2512.06104v1: ARC-AGI Without Pretraining
arc agi2512withoutpretraining
https://velog.io/@sangwu99/Pix2Struct-Screenshot-Parsing-as-Pretraining-for-Visual-Language-Understanding-ICML-2023
Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding (ICML 2023)
visual languagescreenshotparsingpretraining
https://arxiv.org/html/2604.05215v1
Hierarchical Mesh Transformers with Topology-Guided Pretraining for Morphometric Analysis of Brain...
https://arxiv.org/html/2410.05612v2
A Bayesian Model Selection Criterion for Selecting Pretraining Checkpoints
bayesian model selectioncriterionselectingpretrainingcheckpoints
https://openreview.net/forum?id=sqBIm0Irju7
Surprisingly Simple Semi-Supervised Domain Adaptation with Pretraining and Consistency | OpenReview
A strong method for semi-supervised domain adaptation that unlike most recent state of the art approaches does not rely on explicit domain alignment
domain adaptationsurprisinglysimplesemisupervised
https://arxiv.org/abs/2506.22049
[2506.22049] GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation...
Abstract page for arXiv paper 2506.22049: GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling
https://openreview.net/forum?id=7qfkImn0dL
ExPT: Synthetic Pretraining for Few-Shot Experimental Design | OpenReview
Experimental design is a fundamental problem in many science and engineering fields. In this problem, sample efficiency is crucial due to the time, money, and...
few shotexperimental designsyntheticpretrainingopenreview
https://oecd.ai/en/catalogue/metric-use-cases/mtp-advancing-remote-sensing-foundation-model-via-multi-task-pretraining-3
MTP: Advancing Remote Sensing Foundation Model via Multi-Task Pretraining - OECD.AI
We propose a novel model-selection method for dynamic real-life networks. Our approach involves training a classifier on a large body of synthetic network...
remote sensingfoundation model
https://openreview.net/forum?id=hLGJ1qZPdu
On the Generalization Ability of Next-Token-Prediction Pretraining | OpenReview
Large language models (LLMs) have demonstrated remarkable potential in handling natural language processing (NLP) tasks and beyond. LLMs usually can be...
on thegeneralizationabilitynexttoken
https://allenai.org/blog/dolma-3-trillion-tokens-open-llm-corpus-9a0ff4b8da64
Ai2 Dolma: 3 trillion token open corpus for language model pretraining | Ai2
We introduce Dolma, an open dataset from web content, academic publications, code, books, and encyclopedic materials.
for languageai2dolma3trillion
https://arxiv.org/abs/2310.03291v1
[2310.03291v1] SimVLG: Simple and Efficient Pretraining of Visual Language Generative Models
Abstract page for arXiv paper 2310.03291v1: SimVLG: Simple and Efficient Pretraining of Visual Language Generative Models
https://allenai.org/blog/datadecide
DataDecide: How to predict best pretraining data with small experiments | Ai2
Explore the secrets of how language model developers make decisions with DataDecide.
how topredictbest
https://openreview.net/forum?id=ZeFMtRBy4Z
REVE: A Foundation Model for EEG - Adapting to Any Setup with Large-Scale Pretraining on 25,000...
Foundation models have transformed AI by reducing reliance on task-specific data through large-scale pretraining. While successful in language and vision,...
https://openreview.net/forum?id=XlAbMZu4Bo&referrer=%5Bthe%20profile%20of%20Wenhan%20Xiong%5D(%2Fprofile%3Fid%3D~Wenhan_Xiong1)
Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context Length | OpenReview
The quadratic complexity and weak length extrapolation of Transformers limits their ability to scale to long sequences, and while sub-quadratic solutions like...
context lengthmegalodonefficientllmpretraining
https://deepai.org/publication/chinese-clip-contrastive-vision-language-pretraining-in-chinese
Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese | DeepAI
Nov 2, 2022 - 11/02/22 - The tremendous success of CLIP (Radford et al., 2021) has promoted the research and application of contrastive learning for vision...
vision languagechineseclipcontrastivepretraining
https://www.sydney.edu.au/scholarships/b/vision-language-pretraining-model.html
Postgraduate Research Scholarship in Vision Language Pretraining (VLP) model - Scholarships
postgraduate researchin visionscholarshiplanguagepretraining
https://openreview.net/forum?id=QOcukxDfTq
DynaMo: In-Domain Dynamics Pretraining for Visuo-Motor Control | OpenReview
Imitation learning has proven to be a powerful tool for training complex visuo-motor policies. However, current methods often require hundreds to thousands of...
in domainvisuo motordynamodynamicspretraining
https://openreview.net/forum?id=TYdzj1EvBP
How Do Large Language Models Acquire Factual Knowledge During Pretraining? | OpenReview
Despite the recent observation that large language models (LLMs) can store substantial factual knowledge, there is a limited understanding of the mechanisms of...
large language modelshow dofactual knowledge
https://deepai.org/publication/protein-representation-learning-by-geometric-structure-pretraining
Protein Representation Learning by Geometric Structure Pretraining | DeepAI
Mar 11, 2022 - 03/11/22 - Learning effective protein representations is critical in a variety of tasks in biology such as predicting protein function or str...
representation learninggeometric structureproteinpretrainingdeepai
https://openreview.net/forum?id=YhWFvZcahs&referrer=%5Bthe%20profile%20of%20Fabian%20Caba%20Heilbron%5D(%2Fprofile%3Fid%3D~Fabian_Caba_Heilbron3)
Scaling Up Video Summarization Pretraining with Large Language Models | OpenReview
Long-form video content constitutes a significant portion of internet traffic, making automated video summarization an essential research problem. However,...
large language modelsscaling upvideo summarizationpretrainingopenreview
https://openreview.net/forum?id=sV4Y6i1lvE&referrer=%5Bthe%20profile%20of%20Jianwei%20Yang%5D(%2Fprofile%3Fid%3D~Jianwei_Yang1)
Latent Action Pretraining From Videos | OpenReview
We introduce Latent Action Pretraining for general Action models (LAPA), the first unsupervised method for pretraining Vision-Language-Action (VLA) models...
latentactionpretrainingvideosopenreview
https://aclanthology.org/2020.aacl-main.24/
UnihanLM: Coarse-to-Fine Chinese-Japanese Language Model Pretraining with the Unihan Database - ACL...
Canwen Xu, Tao Ge, Chenliang Li, Furu Wei. Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and...
https://www.amazon.science/publications/multitask-pretraining-with-structured-knowledge-for-text-to-sql-generation
Multitask pretraining with structured knowledge for text-to-SQL generation - Amazon Science
Many machine learning-based low-code or no-code applications involve generating code that interacts with structured knowledge. For example, one of the most...
text to sql
https://deepai.org/publication/structure-grounded-pretraining-for-text-to-sql
Structure-Grounded Pretraining for Text-to-SQL | DeepAI
Oct 24, 2020 - 10/24/20 - Learning to capture text-table alignment is essential for table related tasks like text-to-SQL. The model needs to correctly recog...
text to sqlstructuregroundedpretrainingdeepai
https://www.amazon.science/blog/building-geospatial-foundation-models-via-continual-pretraining
Building geospatial foundation models via continual pretraining - Amazon Science
Mar 26, 2024 - New approach enables sustainable machine learning for remote-sensing applications.
foundation modelsbuildinggeospatialviacontinual
https://openreview.net/forum?id=CjTHVo1dvR
Molecular Geometry Pretraining with SE(3)-Invariant Denoising Distance Matching | OpenReview
We propose GeoSSL, a self-supervised learning method using the denoising distance matching for molecular goemetry pretraining.
molecular geometrypretrainingse
https://openreview.net/forum?id=ukgVYJBkvO
Learning the Neighborhood: Contrast-Free Self-Supervised Molecular Graph Pretraining | OpenReview
High-quality molecular representations are essential for property prediction and molecular design, yet large labeled datasets remain scarce. Self-supervised...
the neighborhoodself supervisedmolecular graphlearningcontrast
https://aclanthology.org/2023.unlp-1.4/
GPT-2 Metadata Pretraining Towards Instruction Finetuning for Ukrainian - ACL Anthology
Volodymyr Kyrylov, Dmytro Chaplynskyi. Proceedings of the Second Ukrainian Natural Language Processing Workshop (UNLP). 2023.
gpt 2metadatapretrainingtowards
https://aclanthology.org/2024.knowllm-1.9/
Knowledge Acquisition through Continued Pretraining is Difficult: A Case Study on r/AskHistorians -...
Jan Hoffbauer, Sylwester Sawicki, Marc Ulrich, Tolga Buz, Konstantin Dobler, Moritz Schneider, Gerard De Melo. Proceedings of the 1st Workshop on Towards...
a case study
https://openreview.net/forum?id=1Iuw1jcIrf
MathCoder2: Better Math Reasoning from Continued Pretraining on Model-translated Mathematical Code...
Code has been shown to be effective in enhancing the mathematical reasoning abilities of large language models due to its precision and accuracy. Previous...
https://arxiv.org/abs/2505.02009
[2505.02009] Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale...
Abstract page for arXiv paper 2505.02009: Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs
https://deepai.org/publication/why-is-public-pretraining-necessary-for-private-model-training
Why Is Public Pretraining Necessary for Private Model Training? | DeepAI
Feb 19, 2023 - 02/19/23 - In the privacy-utility tradeoff of a model trained on benchmark language and vision tasks, remarkable improvements have been widel...
for privatemodel trainingpublicpretrainingnecessary
https://openreview.net/forum?id=B8BXHrshMi
FlexRibbon: Joint Sequence and Structure Pretraining for Protein Modeling | OpenReview
Protein foundation models have advanced rapidly, with most approaches falling into two dominant paradigms. Sequence-based language models (e.g., ESM-2) capture...
jointsequencestructurepretrainingprotein
https://openreview.net/forum?id=VYOe2eBQeh
Latent Action Pretraining from Videos | OpenReview
We introduce Latent Action Pretraining for general Action models (LAPA), the first unsupervised method for pretraining Vision-Language-Action (VLA) models...
latentactionpretrainingvideosopenreview
https://openreview.net/forum?id=1mjsP8RYAw
Unsupervised Pretraining for Fact Verification by Language Model Distillation | OpenReview
Fact verification aims to verify a claim using evidence from a trustworthy knowledge base. To address this challenge, algorithms must produce features for...
by languagemodel distillationunsupervisedpretrainingfact
https://aclanthology.org/2025.babylm-main.37/
Pretraining Language Models with LoRA and Artificial Languages - ACL Anthology
Nalin Kumar, Mateusz Lango, Ondrej Dusek. Proceedings of the First BabyLM Workshop. 2025.
language modelsartificial languagespretrainingloraacl
https://deepai.org/publication/on-the-effective-use-of-pretraining-for-natural-language-inference
On the Effective Use of Pretraining for Natural Language Inference | DeepAI
Oct 5, 2017 - 10/05/17 - Neural networks have excelled at many NLP tasks, but there remain open questions about the performance of pretrained distributed w...
natural language inferenceon theeffective use
https://huggingface.co/papers/2406.02214
Paper page - SLTrain: a sparse plus low-rank approach for parameter and memory efficient pretraining
Join the discussion on this paper page
https://openreview.net/forum?id=zBBmV-i84Go
Addressing Resource Scarcity across Sign Languages with Multilingual Pretraining and...
We release the largest available pretraining dataset for sign language across multiple languages and show how multilingual fine-tuning using a unified...
resource scarcitysign languagesaddressingacrossmultilingual