Image to 3D
Turn a photo, sketch or render into a 3D model at /design — accepted formats, single image versus multi-view, what photographs well, and what it costs.
Image to 3D reconstructs a model from a picture instead of a description. It is the same studio at /design — switch the input tab from Text to Image, Image + Text or Multi-view, add your files, and generate. Everything after that is identical to a text prompt: the mesh is repaired, measured, normalised to a printable size and quoted.
Use it when you already know what the thing looks like: a product photo, a concept render, a hand-drawn sketch, a figurine on a table. It is the better path whenever describing the shape in words would take longer than showing it.
Accepted images#
| Limit | |
|---|---|
| Formats | PNG, JPEG, WEBP |
| Size | 10MB per image |
| Single image | Exactly one file |
| Multi-view | 2 to 4 files |
The uploader filters by type in the browser, and the server checks type and size again before a credit is spent. A wrong type is rejected with "unsupported image type", an oversized single image with "image exceeds 10MB limit" (a multi-view set says "an image exceeds the 10MB limit"), and a multi-view set outside the range with "multiview needs 2 to 4 images". None of those cost you anything.
Three ways in#
| Mode | Takes | Prompt | Use it for |
|---|---|---|---|
| Image | One image | Optional | Reconstructing what is in the picture, as it is |
| Image + Text | One image and a prompt | Required | Steering the result — "make it hexagonal and taller" |
| Multi-view | 2 to 4 images | Optional | Objects with a back and sides you care about |
The three differ in more than the file count. In Image mode no prompt reaches the generator at all — the text you type (or the filename, if you leave it blank) names the generation in your history and captions the mesh review, but the reconstruction is made from the picture alone. In Image + Text the prompt is required, is rewritten by the prompt agent, and is sent to the generator alongside the image, so it steers the reconstruction rather than replacing it.
Multi-view sends every view together in one request and is pinned to Tripo — it is the only engine here that reconstructs from several images, so engine routing is skipped for it. Prompts are optional and, when absent, the generation is recorded as "Multiview from N photos".
Taking the photo#
- 1One subject, nothing else
Put the object on a plain surface with nothing behind it. A second object in frame gets reconstructed into the same mesh, and separating them afterwards is more work than reshooting.
- 2Contrast against the background
A white part on a white table has no silhouette to read. Use a background the object clearly is not — a dark cloth for pale objects, a light sheet of paper for dark ones.
- 3Flat, even light
Overcast daylight or a diffused lamp. Hard shadows get read as geometry, and a blown-out highlight erases the surface underneath it. Avoid direct flash.
- 4Shoot slightly above, straight on
A three-quarter view a little above eye level shows the top and one side at once, which gives the reconstruction more to work with than a dead-flat elevation.
- 5Fill the frame, keep it sharp
Get close enough that the object fills most of the picture, and check focus before you upload. Motion blur turns into soft, smeared geometry.
- 6For multi-view, move around the object
Front and back first, then the two sides if you have them. Keep the lighting and the distance roughly the same between shots.
What reconstructs well, and what does not#
| Works | Struggles |
|---|---|
| Single solid objects with a clear outline | Scenes, or two objects in one frame |
| Figurines, toys, ornaments, characters | Wireframe, mesh and lattice structures |
| Product shots on a plain background | Glass, chrome and mirror finishes |
| Chunky mechanical shapes — housings, brackets, knobs | Anything with a specified fit or thread |
| Concept renders and clean 3D turntables | Photos with heavy shadow or a busy background |
| Line-art sketches of a single object | Hair, fur, fabric folds and other fine detail |
| Symmetrical objects seen from the front | Deeply concave shapes hidden from every camera |
The pattern is simple: the generator can only model what a camera could see. A shape that hides its own interior — a deep bowl photographed from the side, the inside of a bottle, the underside of a base — comes back guessed. Multi-view helps, but it does not see through walls.
Credits and time#
| Run | Credits | Typical time |
|---|---|---|
| Image, Fast Draft | 2 | 45–90 seconds |
| Image, Hi-Fi | 3 | 45–90 seconds |
| Image + Text, Fast Draft | 2 | 45–90 seconds |
| Image + Text, Hi-Fi | 3 | 45–90 seconds |
| Multi-view, Fast Draft | 2 | Several minutes — it runs on Tripo |
| Multi-view, Hi-Fi | 3 | Several minutes — it runs on Tripo |
Image and Image + Text run the full prompt-and-routing pipeline whatever mode you pick, so a Fast Draft image is not the shortcut a Fast Draft text prompt is — what draft buys you on the image path is the lower base cost, not much of the wait. In exchange, an image run always reaches an engine that can texture, so it returns PBR maps when textures are on. A text draft cannot.
Credits are charged before the generator is called and refunded automatically if the run errors. Generation is limited to 20 runs an hour per account. Full detail is on /docs/studio/credits.
What you get back#
The same result as a text generation: a repaired mesh normalised to half a Bambu H2S plate (170 × 160 × 170mm), a mesh review with the measured geometry behind it — watertightness, volume, open boundary edges, non-manifold edges, triangle count, bounding box, thinnest dimension, steepest overhang — and a print estimate for one PLA unit at standard quality. Downloads are STL, 3MF and OBJ.
A model whose review comes back as needs redesign cannot be sent to checkout. Reconstructions from a single image fail that check more often than text generations do, usually because a shape the camera never saw came back open. When it happens, the cheapest fix is another photo from a different angle and a multi-view run, rather than the same picture again.