
In a new study, Apple researchers describe SimpleDesign, an improved AI model that can co-create the sequence and structure of proteins. Here are the details.
A little context
Last September, Apple researchers published a study titled “SimpleFold: Folding Proteins is Simpler than You Think,” which detailed a powerful method for predicting the 3D structure of a protein from its amino acid sequence.

Briefly, SimpleFold uses a flow model to generate the 3D structure of a protein directly from its amino acid sequence.
We have explained flow matching in more detail here, but its drawback is that the technique starts with a noiseless background, and learns a relatively straight path to the final result. That contrasts with diffusion models, which typically work by iteratively eliminating noise until they reach a final result.
Both techniques are most commonly (or at least historically) related to image creation, although researchers (including those at Apple) have also explored diffusion models for text and code.
Back to SimpleFold, Apple has basically paired the flow with a common Transformer block (commonly used in text generation), allowing the model to avoid some of the more expensive computational techniques typically used by protein folding models, such as DeepMind’s famous AlphaFold.
Now, Apple researchers have revealed Simple designwhich applies the same push toward simple general-purpose architecture explored with SimpleFold to the broader problem of protein design, rather than simply predicting the 3D structure of proteins.
Simple design
As Apple’s researchers explained in a A new study Title “Simple design: a shared model for protein sequence and structural coding”:
Existing models often rely on a multi-stage training process where autoencoders that tokenize data into latent representations are trained in the first stage. Second, modeling is trained on the latent representation of the autoencoder (s), ie generative modeling in a latent space. We assume that this multi-step training is not necessary to obtain an effective co-design model and therefore present SimpleDesign, an efficient multi-protein design model that is trained directly in the data space.
In other words, while many existing protein design models are based on multi-step processes, SimpleDesign learns to continuously generate amino acid sequences and 3D structures in a single training process.

Many current protein design models work as follows: First, they train separate models to convert protein structures into separate representations, or “tokens.” They then train production models to work with those agents to create new protein sequences and structures.
SimpleDesign skips that intermediate step, learning directly from paired amino acid sequences and 3D coordinates instead of first compressing protein structures into separate tokenized representations.

The way Apple researchers train SimpleDesign is interesting.
They started with more than 2 million pairs of protein sequences and structures, taken primarily from the AFESM dataset, which included structures predicted from the AlphaFold database and additional examples.
During training, they destroyed both parts of each pair: while the amino acids in the sequence were randomly hidden behind masked tokens, the corresponding 3D structure also had noise added to it.
Researchers will also differ in the extent to which each aspect is exploited. If the sequence is largely intact but the structure is severely damaged, the task is similar to protein folding, and the model must recover the structure from the known sequence.
On the other hand, if the structure is mostly intact but the sequence is heavily obscured, it is similar to reverse folding. Then, the model must create a sequence that can produce the defined structure.
And when the two were partially scrambled, the model learned to work on two problems simultaneously, effectively training it for the joint design of proteins.
According to the study, SimpleDesign delivered competitive results across protein co-design, structure generation, and sequence generation benchmarks, despite using a very simple training pipeline.

The researchers also found that SimpleDesign can generate reliable protein structures, and the amino acid sequences it produces are generally as good or better than those produced by the most competitive multimodal models.
Finally, the researchers noted that the results of SimpleDesign are still limited to computer evaluations, because the proteins created have not been tested to confirm that they will fold, work or act safely in real biological systems.
However, the results are quite good, and the whole study (which naturally has more depth about SimpleDesign architecture, training process, standards and results) is worth a good look.
To read the full study, Follow this link.
Worth checking out on Amazon



