X3DStudios

Image to 3D

Turn a photo, sketch or render into a 3D model at /design — accepted formats, single image versus multi-view, what photographs well, and what it costs.

Image to 3D reconstructs a model from a picture instead of a description. It is the same studio at /design — switch the input tab from Text to Image, Image + Text or Multi-view, add your files, and generate. Everything after that is identical to a text prompt: the mesh is repaired, measured, normalised to a printable size and quoted.

Use it when you already know what the thing looks like: a product photo, a concept render, a hand-drawn sketch, a figurine on a table. It is the better path whenever describing the shape in words would take longer than showing it.

Accepted images#

Limit
FormatsPNG, JPEG, WEBP
Size10MB per image
Single imageExactly one file
Multi-view2 to 4 files

The uploader filters by type in the browser, and the server checks type and size again before a credit is spent. A wrong type is rejected with "unsupported image type", an oversized single image with "image exceeds 10MB limit" (a multi-view set says "an image exceeds the 10MB limit"), and a multi-view set outside the range with "multiview needs 2 to 4 images". None of those cost you anything.

Resolution is not the limiting factor
A sharp 1500px photo of a well-lit object beats a 6000px photo of a cluttered shelf. Spend the effort on the subject, not the sensor.

Three ways in#

ModeTakesPromptUse it for
ImageOne imageOptionalReconstructing what is in the picture, as it is
Image + TextOne image and a promptRequiredSteering the result — "make it hexagonal and taller"
Multi-view2 to 4 imagesOptionalObjects with a back and sides you care about

The three differ in more than the file count. In Image mode no prompt reaches the generator at all — the text you type (or the filename, if you leave it blank) names the generation in your history and captions the mesh review, but the reconstruction is made from the picture alone. In Image + Text the prompt is required, is rewritten by the prompt agent, and is sent to the generator alongside the image, so it steers the reconstruction rather than replacing it.

Multi-view sends every view together in one request and is pinned to Tripo — it is the only engine here that reconstructs from several images, so engine routing is skipped for it. Prompts are optional and, when absent, the generation is recorded as "Multiview from N photos".

Multi-view is not photogrammetry
It does not need overlapping frames, a turntable or a calibrated rig. Two to four clean views of the same object is the whole input. Front and back is worth more than four angles ten degrees apart.

Taking the photo#

  1. 1
    One subject, nothing else

    Put the object on a plain surface with nothing behind it. A second object in frame gets reconstructed into the same mesh, and separating them afterwards is more work than reshooting.

  2. 2
    Contrast against the background

    A white part on a white table has no silhouette to read. Use a background the object clearly is not — a dark cloth for pale objects, a light sheet of paper for dark ones.

  3. 3
    Flat, even light

    Overcast daylight or a diffused lamp. Hard shadows get read as geometry, and a blown-out highlight erases the surface underneath it. Avoid direct flash.

  4. 4
    Shoot slightly above, straight on

    A three-quarter view a little above eye level shows the top and one side at once, which gives the reconstruction more to work with than a dead-flat elevation.

  5. 5
    Fill the frame, keep it sharp

    Get close enough that the object fills most of the picture, and check focus before you upload. Motion blur turns into soft, smeared geometry.

  6. 6
    For multi-view, move around the object

    Front and back first, then the two sides if you have them. Keep the lighting and the distance roughly the same between shots.

What reconstructs well, and what does not#

WorksStruggles
Single solid objects with a clear outlineScenes, or two objects in one frame
Figurines, toys, ornaments, charactersWireframe, mesh and lattice structures
Product shots on a plain backgroundGlass, chrome and mirror finishes
Chunky mechanical shapes — housings, brackets, knobsAnything with a specified fit or thread
Concept renders and clean 3D turntablesPhotos with heavy shadow or a busy background
Line-art sketches of a single objectHair, fur, fabric folds and other fine detail
Symmetrical objects seen from the frontDeeply concave shapes hidden from every camera

The pattern is simple: the generator can only model what a camera could see. A shape that hides its own interior — a deep bowl photographed from the side, the inside of a bottle, the underside of a base — comes back guessed. Multi-view helps, but it does not see through walls.

The background-removal toggle currently does nothing
Remove background, and the Full 3D / Silhouette switch beside it, are read from the form and then never placed in any generator payload. They are inert today. Compose the shot as if there were no cut-out step, because there is not one.

Credits and time#

RunCreditsTypical time
Image, Fast Draft245–90 seconds
Image, Hi-Fi345–90 seconds
Image + Text, Fast Draft245–90 seconds
Image + Text, Hi-Fi345–90 seconds
Multi-view, Fast Draft2Several minutes — it runs on Tripo
Multi-view, Hi-Fi3Several minutes — it runs on Tripo
Add 1 for High mesh quality or 4k textures, 2 for Ultra, 8k or a refinement pass.

Image and Image + Text run the full prompt-and-routing pipeline whatever mode you pick, so a Fast Draft image is not the shortcut a Fast Draft text prompt is — what draft buys you on the image path is the lower base cost, not much of the wait. In exchange, an image run always reaches an engine that can texture, so it returns PBR maps when textures are on. A text draft cannot.

Credits are charged before the generator is called and refunded automatically if the run errors. Generation is limited to 20 runs an hour per account. Full detail is on /docs/studio/credits.

What you get back#

The same result as a text generation: a repaired mesh normalised to half a Bambu H2S plate (170 × 160 × 170mm), a mesh review with the measured geometry behind it — watertightness, volume, open boundary edges, non-manifold edges, triangle count, bounding box, thinnest dimension, steepest overhang — and a print estimate for one PLA unit at standard quality. Downloads are STL, 3MF and OBJ.

A model whose review comes back as needs redesign cannot be sent to checkout. Reconstructions from a single image fail that check more often than text generations do, usually because a shape the camera never saw came back open. When it happens, the cheapest fix is another photo from a different angle and a multi-view run, rather than the same picture again.

Check the mesh before you order
A reconstruction can be watertight, well-oriented and still wrong — a smoothed-over hole, a flattened underside, a detail the light hid. The review scores geometry, not resemblance. Rotate the model in the viewer first.