Fashion image editing is a crucial tool for designers to convey their creative ideas by visualizing design concepts interactively. However, current fashion image editing techniques often struggle to accurately identify editing regions and preserve the desired garment texture detail. To address these challenges, we present Detail-Preserved Diffusion Models (DPDEdit), a new multimodal fashion image editing architecture based on latent diffusion models. To precisely locate the editing region, we introduce Grounded-SAM to predict the editing region. To transfer the detail of the given garment texture into the target image, we propose a texture injection and refinement mechanism. This mechanism employs a decoupled cross-attention layer to integrate textual descriptions and texture images, and incorporates an auxiliary U-Net to preserve the high-frequency details of generated garment texture. Additionally, we extend the VITON-HD dataset using a multimodal large language model to generate paired samples with texture images and textual descriptions. Extensive experiments show that our DPDEdit outperforms state-of-the-art methods in terms of image fidelity and coherence with the given multimodal input.