Hard prompts are limited and hard to optimize for large language models

On the Capability and Limitation of Hard Prompt

Machine Learning

Summary

Hard prompts are exact word sequences used to guide language models instead of adjusting their internal settings. The authors show it’s very difficult to find good hard prompts because the problem is computationally complex. They found that short or long hard prompts have intrinsic problems: short ones don’t help much, and long ones often force the model to give the same answer regardless of the question. They also provide rules about when hard prompts can work well across different tasks, offering the first theoretical understanding of how and when hard prompts generalize.

What this means in practice

  • For llm developers: Avoid investing in exhaustive search for hard prompts because finding optimal hard prompts is computationally infeasible.
  • For ai system designers: Design prompting methods that use linear hard prompts to overcome limitations of short or long discrete prompts in transformers.

A theory result. No direct application yet.

Authors

Lijia Yu, Shuaitong Liu, Gaojie Jin, Xinyu Li, Xiao-Shan Gao

Abstract

Prompt engineering has become an indispensable tool for using large language models (LLMs), turning LLMs into task-specific experts without changing their weights. Despite notable theoretical advances in prompt engineering, the theory for the more practical hard or discrete prompts is largely open. In this paper, we try to fill this gap either by providing a complete solution or by making substantial progress on the three core theoretical questions regarding hard prompts. First, we show that determining the existence of a hard prompt for a transformer to solve a downstream task is NP-complete and that finding an optimal hard prompt is NP-hard, which is the first computational complexity result for hard prompting, as far as we know. Second, we show that, unlike soft or continuous prompts, hard prompts have essential limitations: hard prompts are not complete; short hard prompts do not significantly enhance the ability of transformers; and long hard prompts exhibit the "prompt dominating answer phenomenon," meaning that, with high probability, the same answer is given for all queries of the same length. On the other hand, linear hard prompts do not have the limitations of short or long prompts. Third, we provide a tight bound on the size of the task in terms of the prompt length for the performance of prompts on the finite task to generalize to the entire data distribution, leading to a necessary and sufficient condition for generalizability. This is the first result on generalization for prompting, as far as we know. Our findings not only offer the first theoretical insights into hard prompts but also provide provably reliable practical guidance for real-world LLM usage.