Optimizing LAMMPS for dipole interaction in large system

Hello, I am simulating a system of 200 thousand magnetic particles with LAMMPS version 2025.7.22.2 stable_22Jul2025_update2-modified. I have tried different no of CPU cores combination (60, 80, 100, 140) to get the highest speed. Also i have tried to use GPU to accelerate the simulation but it is running very slow. I am not able to find any way to accelerate it, I would be happy to take your suggestions to correct my input script. Please tell if there any other way to calculate long range dipole interaction efficiently. Let me know if you need any other information from my side , my input script is -

dimension 3
boundary p p p
atom_style hybrid sphere dipole
units lj
region Sregion block 0 79.37 0 79.37 0 79.37 units box
create_box 1 Sregion

create_atoms 1 random 200000 807445 Sregion
mass 1 1

pair_style lj/cut/dipole/long 2.5 5
pair_coeff * * 1 1

set type 1 charge 0
set group all dipole/random 2345 2.5

kspace_style ewald/dipole 1.0e-3
neighbor 2.0 bin
neigh_modify delay 1 every 1 check yes
minimize 1.0e-6 1.0e-8 1000 1000
timestep 0.002

velocity all create 5.0 5667 dist gaussian mom yes rot yes

fix 2 all nvt/sphere temp 5 5 0.2 update dipole
thermo_style custom step time temp pe ke etotal press
thermo 1000
run 5000
unfix 2
reset_timestep 0

fix 3 all nvt/sphere temp 1.15 1.15 0.2 update dipole
thermo_style custom step time temp pe ke etotal press
thermo 1000

dump dump_2 all custom 1 sort/outmc* id type x y z mux muy muz
dump_modify dump_2 append no every 500 sort id format line “%d %d %20.15g %20.15g %20.15g %20.15g %20.15g %20.15g”

restart 500 sm_low.dat
run 500000
undump dump_2
unfix 3

Thank you….

  • What is the output of lmp -h?
  • What is your specific command line?
  • What does the first and last part of your output look like (before and after the thermo output)?

This is the main suspect. Plain Ewald summation scales badly with system size even with such a “criminally low” convergence setting.

  1. This is the output of lmp -h …since this version of LAMMPS is without GPU.

Large-scale Atomic/Molecular Massively Parallel Simulator - 22 Jul 2025 - Update 2
Git info (stable / stable_22Jul2025_update2-modified)

Usage example: /home/physics/phd/phz248285/SCRATCH/Stockmayer2d/2D_SM/2D_SMfluid/LAMMPS_dir/build/lmp -var t 300 -echo screen -in in.alloy

List of command-line options supported by this LAMMPS executable:

-echo none/screen/log/both : echoing of input script (-e)
-help : print this help message (-h)
-in none/filename : read input from file or stdin (default) (-i)
-kokkos on/off … : turn KOKKOS mode on or off (-k)
-log none/filename : where to send log output (-l)
-mdi ‘’ : pass flags to the MolSSI Driver Interface
-mpicolor color : which exe in a multi-exe mpirun cmd (-m)
-cite : select citation reminder style (-c)
-nocite : disable citation reminder (-nc)
-nonbuf : disable screen/logfile buffering (-nb)
-package style … : invoke package command (-pk)
-partition size1 size2 … : assign partition sizes (-p)
-plog basename : basename for partition logs (-pl)
-pscreen basename : basename for partition screens (-ps)
-restart2data rfile dfile … : convert restart to data file (-r2data)
-restart2dump rfile dgroup dstyle dfile …
: convert restart to dump file (-r2dump)
-restart2info rfile : print info about restart rfile (-r2info)
-reorder topology-specs : processor reordering (-r)
-screen none/filename : where to send screen output (-sc)
-skiprun : skip loops in run and minimize (-sr)
-suffix gpu/intel/kk/opt/omp: style suffix to apply (-sf)
-var varname value : set index style variable (-v)

OS: Linux “Red Hat Enterprise Linux 9.6 (Plow)” 5.14.0-570.12.1.el9_6.x86_64 x86_64

Compiler: GNU C++ 11.2.0 with OpenMP 4.5
C++ standard: C++17
Embedded fmt library version: 10.2.0
Embedded JSON class version: 3.12.0

MPI v3.1: Open MPI v4.1.6, package: Open MPI sksharma.vfaculty@vsky008 Distribution, ident: 4.1.6, repo rev: v4.1.6, Sep 30, 2023

Accelerator configuration:

FFT information:

FFT precision = double
FFT engine = mpiFFT
FFT library = KISS

Active compile time flags:

-DLAMMPS_GZIP
-DLAMMPS_PNG
-DLAMMPS_JPEG
-DLAMMPS_SMALLBIG
sizeof(smallint): 32-bit
sizeof(imageint): 32-bit
sizeof(tagint): 32-bit
sizeof(bigint): 64-bit

Available compression formats:

Extension: .gz Command: gzip
Extension: .bz2 Command: bzip2
Extension: .zst Command: zstd
Extension: .xz Command: xz
Extension: .lzma Command: xz
Extension: .lz4 Command: lz4

Installed packages:

DIPOLE KSPACE MANYBODY MOLECULE RIGID

List of individual style options included in this LAMMPS executable

  • Atom styles:

angle atomic body bond charge
dipole ellipsoid full hybrid line
molecular sphere template tri

  • Integrate styles:

respa verlet

  • Minimize styles:

cg fire/old fire hftn quickmin
sd

  • Pair styles:

adp airebo airebo/morse atm bop
born born/coul/long born/coul/msm buck buck/coul/cut
buck/coul/long buck/coul/msm buck/long/coul/long comb
comb3 coul/cut coul/debye coul/dsf coul/long
coul/msm coul/streitz coul/wolf meam/c reax
reax/c mesont/tpm eam eam/alloy eam/cd
eam/cd/old eam/fs eam/he edip edip/multi
eim extep gw gw/zbl
hbond/dreiding/lj hbond/dreiding/morse hybrid
hybrid/omp hybrid/molecular hybrid/molecular/omp
hybrid/overlay hybrid/overlay/omp hybrid/scaled
hybrid/scaled/omp lcbop lj/charmm/coul/charmm
lj/charmm/coul/charmm/implicit lj/charmm/coul/long
lj/charmm/coul/msm lj/charmmfsw/coul/charmmfsh
lj/charmmfsw/coul/long lj/cut lj/cut/coul/cut lj/cut/coul/long
lj/cut/coul/msm lj/cut/dipole/cut lj/cut/dipole/long
lj/cut/tip4p/cut lj/cut/tip4p/long lj/expand
lj/long/coul/long lj/long/dipole/long
lj/long/tip4p/long lj/sf/dipole/sf local/density meam/spline
meam/sw/spline morse nb3b/harmonic nb3b/screened polymorphic
rebo rebomos soft sw sw/angle/table
sw/mod table tersoff tersoff/mod tersoff/mod/c
tersoff/table tersoff/zbl threebody/table tip4p/cut tip4p/long
vashishta vashishta/table yukawa zbl zero

  • Bond styles:

fene fene/expand gromos harmonic hybrid
morse quartic table zero

  • Angle styles:

charmm cosine cosine/squared dipole harmonic
hybrid table zero

  • Dihedral styles:

charmm charmmfsw harmonic hybrid multi/harmonic
opls table zero

  • Improper styles:

cvff harmonic hybrid umbrella zero

  • KSpace styles:

ewald ewald/dipole ewald/dipole/spin ewald/disp
ewald/disp/dipole zero msm msm/cg
pppm pppm/cg pppm/dipole pppm/dipole/spin
pppm/disp pppm/disp/tip4p pppm/stagger pppm/tip4p

  • Fix styles

adapt addforce ave/atom ave/chunk ave/correlate
ave/grid ave/histo ave/histo/weight ave/time
aveforce balance box/relax cmap deform
deposit ave/spatial ave/spatial/sphere lb/pc
lb/rigid/pc/sphere reax/c/bonds reax/c/species dt/reset
efield ehex enforce2d evaporate external
gravity halt heat indent langevin
lineforce momentum move nph nph/sphere
npt npt/sphere nve nve/limit nve/noforce
nve/sphere nvt nvt/sllod nvt/sphere pair
planeforce press/berendsen press/langevin print property/atom
qeq/comb rattle recenter restrain rigid
rigid/nph rigid/nph/small rigid/npt rigid/npt/small rigid/nve
rigid/nve/small rigid/nvt rigid/nvt/small rigid/small set
setforce shake spring spring/chunk spring/self
store/force store/state temp/berendsen temp/rescale
thermal/conductivity tune/kspace vector viscous
wall/harmonic wall/lj1043 wall/lj126 wall/lj93 wall/morse
wall/reflect wall/region wall/table

  • Compute styles:

aggregate/atom angle angle/local angmom/chunk bond
bond/local centro/atom centroid/stress/atom chunk/atom
chunk/spread/atom cluster/atom cna/atom com
com/chunk coord/atom count/type mesont dihedral
dihedral/local dipole dipole/chunk displace/atom erotate/rigid
erotate/sphere erotate/sphere/atom fragment/atom global/atom
group/group gyration gyration/chunk heat/flux improper
improper/local inertia/chunk ke ke/atom ke/rigid
msd msd/chunk omega/chunk orientorder/atom
pair pair/local pe pe/atom pressure
property/atom property/chunk property/grid property/local rdf
reduce reduce/chunk reduce/region rigid/local slice
stress/atom temp temp/chunk temp/com temp/deform
temp/partial temp/profile temp/ramp temp/region temp/sphere
torque/chunk vacf vacf/chunk vcm/chunk

  • Region styles:

block cone cylinder ellipsoid intersect
plane prism sphere union

  • Dump styles:

atom cfg custom atom/mpiio cfg/mpiio
custom/mpiio xyz/mpiio grid grid/vtk image
local movie xyz

  • Command styles

angle_write balance change_box create_atoms create_bonds
create_box delete_atoms delete_bonds box kim_init
kim_interactions kim_param kim_property kim_query
reset_ids reset_atom_ids reset_mol_ids message server
dihedral_write displace_atoms info minimize read_data
read_dump read_restart replicate rerun run
set velocity write_coeff write_data write_dump
write_restart

  1. Specific command line to start simulaiton is- mpiexec -np 60 --oversubscribe lmp -in in.sm

  2. The output file is- (before thermo)

LAMMPS (22 Jul 2025 - Update 2)
using 1 OpenMP thread(s) per MPI task
WARNING: Atom style hybrid defines both, per-type and per-atom masses; both must be set, but only per-atom masses will be used (src/atom_vec_hybrid.cpp:132)
Created orthogonal box = (0 0 0) to (79.37 79.37 79.37)
4 by 3 by 5 MPI processor grid
Created 200000 atoms
using lattice units in orthogonal box = (0 0 0) to (79.37 79.37 79.37)
create_atoms CPU = 0.228 seconds
Setting atom values …
200000 settings made for charge
Setting atom values …
200000 settings made for dipole/random
Switching to ‘neigh_modify every 1 delay 0 check yes’ setting during minimization
EwaldDipole initialization …
using 12-bit tables for long-range coulomb
G vector (1/distance) = 0.63401133
estimated absolute RMS force accuracy = 0
estimated relative force accuracy = 0
KSpace vectors: actual max1d max3d = 277745 51 546363
kxmax kymax kzmax = 51 51 51
Generated 0 of 0 mixed pair_coeff terms from geometric mixing rule
Neighbor list info …
update: every = 1 steps, delay = 0 steps, check = yes
max neighbors/atom: 2000, page size: 100000
master list distance cutoff = 7
ghost atom cutoff = 7
binsize = 3.5, bins = 23 23 23
1 neighbor lists, perpetual/occasional/extra = 1 0 0
(1) pair lj/cut/dipole/long, perpetual
attributes: half, newton on
pair build: half/bin/atomonly/newton
stencil: half/bin/3d
bin: standard
Setting up cg style minimization …
Unit style : lj
Current step : 0
Per MPI rank memory allocation (min/avg/max) = 152.7 | 179.5 | 236.5 Mbytes
Step Temp E_pair E_mol TotEng Press
0 0 5.124865e+13 0 5.124865e+13 8.1998003e+13
924 0 -15.137323 0 -15.137323 -0.75413359
Loop time of 25986.6 on 60 procs for 924 steps with 200000 atoms

99.5% CPU use with 60 MPI tasks x 1 OpenMP threads

Minimization stats:
Stopping criterion = max force evaluations
Energy initial, next-to-last, final =
51248649889901.5 -15.1364393934177 -15.1373230231984
Force two-norm initial, final = 1.058013e+21 438.76673
Force max component initial, final = 1.3384255e+20 74.731329
Final line search alpha, max atom move = 0.00045913384 0.034311682
Iterations, force evaluations = 924 1000

MPI task timing breakdown:

Section | min time | avg time | max time |%varavg| %total

Pair | 41.72 | 50.51 | 55.955 | 46.4 | 0.19
Kspace | 21484 | 24241 | 25927 | 725.0 | 93.28
Neigh | 0.41331 | 0.53901 | 0.76857 | 14.1 | 0.00
Comm | 2.4255 | 1692 | 4452.6 |2750.3 | 6.51
Output | 0 | 0 | 0 | 0.0 | 0.00
Modify | 0 | 0 | 0 | 0.0 | 0.00
Other | | 2.824 | | | 0.01

Nlocal: 3333.33 ave 3498 max 3123 min
Histogram: 2 1 7 7 4 11 11 6 9 2
Nghost: 12950.7 ave 13368 max 12536 min
Histogram: 3 1 7 6 13 13 7 5 3 2
Neighs: 957389 ave 1.02741e+06 max 863892 min
Histogram: 4 2 5 5 7 7 8 7 10 5

Total # of neighbors = 57443318
Ave neighs/atom = 287.21659
Neighbor list builds = 16
Dangerous builds = 0
EwaldDipole initialization …
using 12-bit tables for long-range coulomb
G vector (1/distance) = 0.63401133
estimated absolute RMS force accuracy = 0
estimated relative force accuracy = 0
KSpace vectors: actual max1d max3d = 277745 51 546363
kxmax kymax kzmax = 51 51 51
Generated 0 of 0 mixed pair_coeff terms from geometric mixing rule
Setting up Verlet run …
Unit style : lj
Current step : 924
Time step : 0.002

And after thermo ….the simulation is not done yet.

So is there any alternative of Ewald summation for my system size. or what can I do to improve it. Thank for your reply.

This is the confirmation: over 90% of the time spent in KSpace.

Have you tried pppm/dipole?

No, I have not tried pppm/dipole. Actually for small system I have taken ewald/dipole, so I was just taking it same for large system. Now I will try ppm/dipole for sure and check the things. But will pppm/dipole change the dynamics of system or it just do the calculation like ewald/dipole? Also how fast it is compare to ewald/dipoe, i mean should I expect very much change in calculation time?

Sorry to ask these questions as I have not explored ppppm/dipole calculations till now, but i will.

Thank you for your help.

This cannot be described easily in absolute terms at a general level.
A regular Ewald summation and a Particle-Particle, Particle-Mesh calculation (PPPM) have very different performance characteristics and scaling with system size. For supporting dipoles, the calculations are more complex, but the general performance trends should be similar, i.e. O(N^(3/2)) for Ewald and O(N log(N)) for PPPM over a wide variety of system sizes for dense systems. You can learn about this in text books on MD simulations where they discuss long-range electrostatics and in the most recent LAMMPS paper. Among widely used MD codes, PME is more commonly used than PPPM, but the performance is in both cases bounded by the speed of the 3d-FFTs.

okay, Thanks a lot for your time.