AI 이미지 생성

AI로 AI를 위한 Prompt 작성하기

사용자가 「Alice가 발코니에서 멍 때리고 있다」고 말합니다. 이미지 생성 모델에게는 길고 정확한 시각적 설명이 필요하며, 그 한 문장만으로는 전혀 부족합니다. 중간에 번역 단계가 있습니다:LLM을 사용해 일상 언어를 이미지 생성 언어로 번역하는 것입니다.

예시 1:발코니에서 멍 때리기
사용자가 말함
「Alice가 발코니에서 멍 때리고 있다」
LLM이 이미지 생성 Prompt로 번역
A young woman named Alice standing on a sunlit balcony, leaning against the railing, gazing into the distance with a dreamy expression. She has shoulder-length dark hair, wearing a white blouse with a small star necklace. Soft afternoon golden hour lighting, potted plants on the balcony, blurred city skyline in background. Illustration style, warm color palette, peaceful mood. Upper body to full body composition.
이미지 생성 모델 출력
발코니에서 멍 때리는 Alice
예시 하나 더
예시 2:아침 요리
사용자가 말함
「Alice가 요리하고 있다」
LLM이 이미지 생성 Prompt로 번역
A young woman named Alice in a bright modern kitchen during morning time, cooking breakfast. She has shoulder-length dark hair, wearing a casual cardigan over a white top with a star necklace. Warm natural light streaming through windows, kitchen utensils and ingredients on counter, steam rising from pan. Illustration style, cozy domestic atmosphere, soft warm tones. Wide shot showing kitchen environment.
이미지 생성 모델 출력
주방에서 아침을 만드는 Alice
왜 이 번역 레이어가 필요한가요?
① 사용자는 이미지 생성 Prompt를 작성할 줄 모릅니다:golden hour lighting을 지정해야 하는지 illustration style을 지정해야 하는지 모릅니다
② 이미지 생성 모델은 모호한 의도를 이해하지 못합니다:「멍 때리다」는 모델에게 시각적 설명이 아닙니다
③ 각 이미지 생성 모델마다 방언이 다릅니다:Midjourney, DALL-E, Stable Diffusion은 각기 다른 Prompt 스타일을 선호합니다
사용자가 상상하는 것과 이미지 생성 모델이 필요로 하는 것은 완전히 다른 언어입니다。Alice에서는 모든 이미지 생성 뒤에 번역을 수행하는 LLM이 있습니다:사용자의 한 문장을 수백 token의 상세한 시각적 설명으로 확장합니다. 이건 부가 기능이 아니라 필수 아키텍처입니다.