Creation of new and analysis of given data for MLP training

Good afternoon to everyone!

I’m new in the ML potentials creation and I have few questions. I really appreciate if anyone may help me.

I’m interested in some details of MLP training datasets creation. It’s obviously that a dataset containing wide range of structures has to be prepared for a training. But how definitely the structures are chosen?

I mean, it’s logical (?) to take a pressure-temperature phase diagram and mark all the transition points and create structures at the points and near the points for the dataset. But how many states from the vicinity of a transition point should be taken? And what is the vicinity - I mean, how big dT and dP near the point it should be? And also the parts inside the transition lines also have to be covered well. So how usually is this done? All is divided with lines of constant pressure and temperature with step, for example, N GPa and M Kelvin, and then states at the cross points of the lines are created? If it is literally the block division then what pressure step and what temperature step are enough? Or is it done somehow different?

Thank you for your help!

Hi Mary, hope you can provide more information about the use of your developed MLP.

Hi

Well, the potential isn’t developed at all yet… I’m trying to solve the problem of dataset creation as I wrote in the start message. But none answers my questions unfortunately.

So the only one thing I may provide now is the well known fact that quality of the dataset powerfully impact on the potential’s quality - I have made one wrong, despite the fact I gathered huge volume of structures, and it totally didn’t work.

So now I ask about details and hope some people who may help me will come here and give the help.

P.S.

If I misunderstood you and you mean what for I want to create a MLP, then the answer is that I would like to model a system’s behaviour under high pressure and temperature conditions