Introduction to Structural Biology
Beate Bersch, researcher at the CNRS, Biomolecular NMR group
Institut de Biologie Structurale,
Tel 04 57 42 85 12, email [Link]@[Link]
Research:
Application of biomolecular NMR spectroscopy to the study of proteins.
Proteins involved in bacterial outer membrane interactions, membrane proteins
Proteins of the mitochondrial protein import machinery
Viral proteins
Structural Biology
The very core of life lies in the interaction between atoms in molecules.
For example, proteins interacting with each other or with DNA/RNA. Lipids
forming membranes holding proteins in place to facilitate transport across cell
membranes, recognition of signalling molecules, receptor activation...
The underlying principle of function is structure
Structural biology aims at the mechanistic understanding of the function of
biomolecules and the underlying cellular processes by the characterisation of
biomolecular structures, interaction, and dynamics
Definition
Structural biology is the study of the molecular structure and dynamics of
biological macromolecules, particularly proteins and nucleic acids, and how
alterations in their structures affect their function.
Structural biology incorporates principles of molecular biology, biochemistry and
biophysics.
[Link] [Link]
[Link]
[Link] (Structural biology, a golden era)
Contributions from structural biology
• protein classification
• structure activity relationship (analysis of molecular interactions, identification
of active sites …)
• rational approach for understanding certain pathologies related to molecular
dysfunction
• rational design of active molecules, modified enzymes …
Quick introduction to protein structure
Proteins are composed of amino acids, connected by peptide bonds in a linear
chain that folds into the more complex protein structure
[Link]
Introduction to protein structure
Amino Acid 3-Letter 1-Letter Side chain S i d e c h a i n Hydropathy
amino acids polarity _ acidity o r index _
basicity _
Alanine Ala A nonpolar neutral 1.8
Arginine Arg R polar basic (strongly) -4.5 A measure of
Asparagine Asn N polar neutral -3.5 polarity of an amino
Aspartic acid Asp D polar acidic -3.5
acid residue; the
free energy of
Cysteine Cys C polar neutral 2.5
transfer of the
Glutamic acid Glu E polar acidic -3.5
residue from a
Glutamine Gln Q polar neutral -3.5
medium of low
Glycine Gly G nonpolar neutral -0.4 dielectric constant
Histidine His H polar basic (weakly) -3.2 to water.
Isoleucine Ile I nonpolar neutral 4.5 The larger the
Leucine Leu L nonpolar neutral 3.8 number is, the more
Lysine Lys K polar basic -3.9 hydrophobic the
Methionine Met M nonpolar neutral 1.9 amino acid.
Phenylalanine Phe F nonpolar neutral 2.8
Proline Pro P nonpolar neutral -1.6
Serine Ser S polar neutral -0.8
Threonine Thr T polar neutral -0.7
Tryptophan Trp W nonpolar neutral -0.9
Tyrosine Tyr Y nonpolar neutral -1.3
Valine Val V nonpolar neutral 4.2
Amino acid properties: graphical representation
Introduction to protein structure
Primary structure
Primary structure (amino acid sequence) can be obtained from the gene sequence
(DNA, cDNA, mRNA …).
Protein folding:
The amino-acid sequence of a protein determines its native conformation. A
protein molecule folds spontaneously during or after biosynthesis on the ribosome.
Folding also depends on the solvent (water or lipid bilayer), the concentration of
salts, the pH, the temperature, the possible presence of cofactors and of molecular
chaperones.
The correct three-dimensional structure is essential to function, although (some
parts of) functional proteins may remain unfolded. Failure to fold into native
structure generally produces inactive proteins, but in some instances misfolded
proteins have modified or toxic functionality. Several neurodegenerative and other
diseases are believed to result from the accumulation of amyloid fibrils formed by
misfolded proteins
[Link]
Introduction to protein structure
Three-dimensional structure of proteins
Hierarchical principle: primary, secondary, tertiary, quaternary structure
Secondary structure
Repetitive local structure of the main chain (backbone)
a helix b sheet
Introduction to protein structure
Tertiary structure
Fold of the protein that forms and is stabilized through long-range contacts of
the protein chain (van-der-Waals, hydrogen bonds, electrostatic interactions).
Many proteins are folded around a hydrophobic core..
Quaternary structure
Assembly of different protein chains in multimeric structures
Protein representation
C-Jun CBM22 GH10 CBM22 Dockerin CE1
MPKVSVIMTSYNKSDYVAKSISSILSQTFSDFELFIMDDNSNET
PI3K LNVIRPFLNDNRVRFYQSDISGVKERTEKTRYAALINQAIEMAE
GEYITYATDDNIYMPDRLLKMVRELDTHPEKAVIYSASKTYHLN
ENRDIVKETVRPAAQVTWNAPCAIDHCSVMHRYSVLEKVKEKF
GSYWDESPAFYRIGDARFFWRVNHFYPFYPLDEELDLNYITDQ
SIHFQLFELEKNEFVRNLPPQRNCRELRESLKKLGMG
Now you've seen some beautiful pictures already …
how to address biomolecular structure ?
Towards the observation of the molecular bases of life
plant animal bacteria poxvirus virus proteins small atoms
ADN
ADN molecules
cells cells ribosome
Meters
Atomic
resolution
How can we “see” macromolecules?
Based on interaction between radiations and matter
The wavelength of the radiation used should be suited to the size of the smallest
details we want to observe:
ØResolution limit in diffraction experiments is λ/2
Proteins and macromolecules assemblies are too small (1 to 50 nm, 10 to 500 Å)
to be observed with visible light (300 nm < l < 800 nm, n ≈ 430–770 THz)
To "see" atoms, a wavelength of about 1 Å is required ! -> X-ray crystallography,
-> electron beam in EM
(NMR is a special case as we will see later)
Principal Techniques in Structural Biology
Electron Small-angle NMR X-ray
Microscopy Scattering crystallography
Limitations or bottle necks:
Low molecular Limited resolution High molecular Crystallization,
mass, mass, solubility phasing
(limited resolution)
Structural Biology and the electromagnetic spectrum
EPR
NMR Raman EM neutrons
micros- X-ray/Saxs
MRI cope
synchrotron
1 eV = 1.6×10-19 J
stability of C-C bonds: ≈ 400 kJ/mol ≈ 3 eV
Atomic resolution structure from protein crystallography
crystallization screen crystals X-ray diffraction
3D structure atomic model electron density map
Basics of X-ray diffraction
A beam of X-rays forms a Scatterers re-radiate a small portion of the beam intensity
spot on a photographic as a spherical wave. If scatterers are arranged symmetri-
film cally, these spherical waves cancel through destructive
interference, or add constructively forming the diffraction
pattern (Braggs Law).
A crystal diffracts the beam in By measuring the positions of the
many directions, showing fainter spots, the arrangement of atoms
spots on the film. in the crystal can be found
[Link]
Atomic resolution structure from NMR
NMR sample NMR data acquisition resonance assignment
structural ensemble
structure calculation structural parameters
Basics of NMR
NMR spectroscopy exploits the interaction of nuclear spins with an external magnetic field.
B0 Larmor frequency:
B0 w=g×B0
w=g×B0
N S
N
S
B1
resonance !
E
B0
DE energy absorption gives rise to the NMR
signal, the FID, that allows measuring the
B0 Larmor frequency
Here, the energy difference between the two Atoms with nuclear spin: 1H, 13C, 15N, 31P
states is less than 0.1 cal/mol, and, with E = Each nuclear spin "feels" its neighbors
hn, is generally expressed as a frequency => interactions through bonds or
(radiofrequencies between 50 MHz and 1 through space that can be measured and
GHz, corresponding to 0.3 m < l < 6 m). h = quantified => structure
6.63×10-34 J.s
Initiating and performing a protein structure determination
Steps towards a three-dimensional protein structure
- Which method for which protein ?
- Protein sequence analysis
composition, sequence homology, secondary structures, topology predictions …
- Getting the protein ready for structure determination
recombinant protein, tags, fusion proteins, protein purification
isotopic labelling (NMR), introduction of selenomethionine (crystallography)
in vitro or cell-free protein synthesis
- Quality control
SDS-PAGE, immunoblotting, mass spectrometry, N-terminal sequencing,
light scattering, analytical ultracentrifugation, circular dichroism, NMR
NMR and X-ray crystallography
Towards structural information at atomic resolution …
X-ray crystallography:
FT
Model
building
NMR
Resonance Structural
assignment parameters Structure calculation
Which method to choose ?
Pros and Cons of X-ray crystallography and NMR spectroscopy …
X-ray:
+ well established, highly automated
+ more direct (mathematical) image construction
+ more objective interpretation of experimental data
+ quality indicators available (R factor, resolution)
+ structure of large molecules and assemblies can be determined
+ produces a single model, easy to visualize and to interpret
- need for well-diffracting crystals (not suited for highly flexible proteins)
- need for heavy atom derivatives or structures of similar protein (phases !)
- unnatural, non-physiological environment
- influence of crystal packing, resolved structure may differ from solution structure
- details of mobile parts are unresolved
Which method to choose ?
Pros and Cons of X-ray crystallography and NMR spectroscopy …
NMR:
+ solution conditions, more physiological, no crystallization artifacts
+ experimental parameters can be varied (pH, temperature, salt concentration …)
+ provides information on dynamics
+ easy identification of ligand binding sites (ligand screening)
+ secondary structures can be determined from limited experimental data
+ good for fold determination of mutants
+ secondary structure propensity in intrinsically denatured proteins (IDPs)
+ protein folding studies
- solubility, stability, need for monodisperse solution
- size limitation, need for isotopic labeling (deuteration) in larger proteins
- lack of established quality indicators (structure validation less straight-forward)
- more subjective data interpretation
- side chain position less well resolved, produces structural ensemble
1st step: Let’s have a look at the protein sequence
Protein sequences are generally known, they can be obtained from a data
bank
The UniProt Knowledgebase (UniProtKB) is the central access point for
extensive curated protein information. It comprises two sections:
UniProtKB/Swiss-Prot which is manually annotated and reviewed
UniProtKB/TrEMBL which is automatically annotated and is not reviewed
The UniProt Reference Clusters (UniRef) databases provide clustered sets
of sequences from the UniProtKB and selected UniProt Archive records to obtain
complete coverage of sequence space at several resolutions while hiding redundant
sequences.
Release 2020_04 of 12-Aug-2020 of
UniProtKB/TrEMBL contains 188 961 949
sequence entries.
(65 378 749 were in the Release 2016_07)
[Link]
Some sequence analysis tools
[Link] Bioinformatics Resource Portal
Expasy is the bioinformatics resource portal of the SIB Swiss
Institute of Bioinformatics (more about its history:
[Link]
It is an extensible and integrative portal which provides
access to over 160 databases and software tools, developed
by SIB Groups and supporting a range of life science and
clinical research domains, from genomics, proteomics and
structural biology, to evolution and phylogeny, systems
biology and medical chemistry.
Protein sequence analysis (still useful)
My favorite tools:
- Protein parameters (MW, pI, atomic composition, extinction coefficient …):
[Link]
- Similarity search: [Link] (options !)
- Use UniRef rather than UniProt (redundancy)
- Protein sequence alignment: [Link] - the expresso module allows
to explicitly enter structural information
- Nice alignment views:
[Link]
- Protein sequence motifs: [Link] (scan sequences
against collection of motifs, create your own motif and scan databases against it)
- Secondary structure prediction: [Link]
- Protein fold recognition, "threading": Protein threading is a method of protein
modeling which is used to model those proteins which have the same fold as
proteins of known structures, but do not have homologous proteins with
known structure. [Link]
Primary sequence analysis
identification of functional sites
cellular location, identification of transmembrane helices …
amino acid composition, molecular mass, pI, absorption coefficient …
secondary structure prediction, identification of low complexity regions,
globularity, stability …
homology modeling, threading
first indications for feasibility of structural study
Example:
448 aa protein fragment from Hepatitis virus (non structural protein 5A)
PredictProtein output:
sequence number
predicted functional sites
predicted helix / sheet
predicted buried / exposed
predicted disordered
predicted disordered
Sequence analysis (1): the N-terminal helix
W L R D IW DW I C EV L S D F KTW L K A K
helical wheel plot
[Link]
Sequence analysis (2): residues 32-202
Predicted secondary structures and alternating buried and exposed regions indicate
a folded protein chain.
the structure was solved by X-ray
crystallography
Sequence analysis (3): residues 191-369
Prediction of Intrinsically Unstructured Proteins
[Link]
Sequence analysis (3): residues 191-369
NMR chemical shift analysis shows presence of transient a-helices:
205-221, 251-266, 292-306
DeepMind, alphafold
• high-quality predictions for the shape of every single protein in the human body, as
well as for the proteins of 20 additional organisms
• curated by EBI-EMBL and accessible in a data base: [Link]
[Link]
[Link]
[Link]
Alphafold and structural biology
now that protein structures can be predicted with high confidence, where do we go ?
new approaches: use the information from the predicted model structures for a
quicker (and more reliable ???) data interpretation (molecular replacement,
automatic NOE assignment)
Sample Preparation (1): Sample requirements
X-ray crystallography:
- the sample corresponds to a protein crystal
- conditions for successful crystallization: several milligram of a pure (≥ 90%)
and homogeneous protein sample in a buffer that promotes sample stability,
homogeneity and monodispersity
- protein should be predicted to be structured (avoid flexible termini)
- Selenomethionine labeling (recombinant protein), introduction of heavy
atoms
NMR
- the sample is a relatively concentrated protein solution (≥ 200 uL, ≥ 0.1 mM)
- The protein sample should be stable, ideally for several days at room
temperature (or even higher !)
- -> homogeneity, stability, solubility are important factors
- Protein can be partly or completely unstructured
- Size limitation: ≤ 40 kDa for a structural study (divide and conquer)
- Isotopic labeling required (15N, 13C, 2H) – recombinant protein
Protein sources
Purification from the native host
- only a few proteins are present in sufficient quantities for a structural study
Example: Cytochrome C-551 from Ectothiorhodospira halophila:
53 mmol (≈400 mg) from 1500 g of cells grown in 500 L of culture medium for 6
days in anaerobic conditions …
-> production of a recombinant protein in an adapted host (Escherichia
coli, Saccharomyces cerevisae, Pichia pastoris, insect cells, CHO …)
- gene amplification and cloning in an expression vector (pET series …)
- transformation of the host with the expression vector
- multiplication of the recombinant host and induction of protein production
- purification
- quality control
- sample preparation and conditioning
The expression system
The gene coding for the recombinant protein is cloned into an expression
vector under the control of an inducible promoter.
pT7 (RNA polymerase of T7 bacteriophage), pBAD (arabinose), lac
(lactose/IPTG) etc …
pET system:
host: E. coli BL21(DE3), strain
deficient in lon and ompT
proteases. The gene coding for
the T7 RNA polymerase is
integrated in the bacterial
chromosome under the control
of a lac promoter.
Vectors: pET (Novagen), many
different vectors proposing
various fusion proteins, tags,
protease cleavage sites,
antibiotic resistance genes etc
Increasing protein yields: common tags and fusion proteins
Example: Comparing different fusion proteins
Protein overexpression: producing large quantities of a "foreign" protein in a short
time = stress for the host
Problem: the expressed protein is not always well folded (impossibility to form
disulfide in the cytoplasm, lack of cofactors, chaperones …) -> the protein
precipitates in inclusion bodies
Aim: get large quantities of soluble protein for structural studies
tiré de: Marblestone et al., (2006) Prot Science 15, 182-189
see also: [Link]
Isotopic labeling, introduction of selenomethionine
Bacteria or yeast cells can be grown in minimal media that do not contain any
amino acids
Isotopic labeling for NMR: the desired isotopes are introduced
by using simple and cheap molecules as sole carbon and/or
nitrogen sources: 13C-glucose, 15NH4Cl.
Bacteria and yeasts can also be grown in heavy water
(deuteration).
Selenomethionine for X-ray crystallography (MAD phasing):
Bacteria: No need for auxotrophic strains as the methionine biosynthesis can be
blocked by high concentrations of isoleucine, lysine and threonine.
Selenomethionine is added to the culture medium as sole source for methionine.
Eucaryotes: are naturally auxotrophic for methionine
In vitro or cell-free protein synthesis
Principle: transcription and translation take place in vitro (cell free)
plasmids bearing a T7 promoter
addition of recombinant T7 RNA
polymerase
the translation machinery is isolated from
E. coli lysates (prepared in the lab or
commercially available kits)
≤ 5 mg of protein
≤ 3x50 mg de protéine
Cell-free protein synthesis (in vitro)
Advantages:
- composition of expression medium is adaptable
- expression of proteins that are toxic for living E. coli cells
- co-expression of several proteins (complex formations …)
- selective labeling by one or more amino acids (scrambling still possible)
- incorporation of selectively labeled amino acids
- less proteases, more stable proteins
- easier purification
(NMR: reaction medium can be analyzed directly w/o purification)
- membrane proteins: protein synthesis in presence of detergents, liposomes …
Disadvantages:
- more expensive than bacterial cultures
- yield
Quality control
1. SDS polyacrylamide gel electrophoresis:
quantity, purity, integrity (approximate size)
2. Immunoblotting:
protein identification (! need for antibodies that recognize protein or tag !)
specific detection of small quantities
Principle
- transfer of proteins from an SDS gel to a nitrocellulose membrane
- reaction with mono- or polyclonal antibodies or other specific proteins
- reaction with a secondary, labeled antibody (HRP, AP …)
- development (colorimetric detection of the target protein)
Quality control
3. N-terminal sequencing
Identification of the protein's N-terminal (cleavage of N-terminal methionine,
N-terminal proteolytic degradation …)
4. Amino acid analysis
Hydrolysis, detection and quantification of individual amino acids:
composition, quantification, experimental extinction coefficient
5. Mass spectrometry:
identification of the protein by its molecular mass (± 1-3 Da), in addition,
detection of mutations, disulfide bonds, oxydation, ligands …
quantification of isotopic labeling, SeMet insertion
determination of sample homogeneity, integrity (proteolytic degradation,
cleavage of N-terminal methionine …)
Quality control: mass spectrometry
Separation of ionized macromolecules according to their mass to charge
ratio :
Principle: 1. ionization and transfer to the gas phase
2. separation of individual ions (space, time) according to m/z
3. quantification of ions at a given m/z
± 1Da
± 5Da
Contrôle de qualité: spectrométrie de masse
54009.94 Da 30606.69 Da
top: ESI-MS mass spectra of two different
proteins, insets represent the molecular
mass spectra obtained after deconvolution
of the raw data, theoretical masses are given
in red.
right: ESI mass spectra of a recombinant
protein titrated with metal ions.
Quality control
6. Light scattering, SEC-MALS and/or analytical ultracentrifugation
Information on the molecular size -> oligomerization, aggregation, relative populations
(monodisperse/polydisperse solution)
see also: [Link]
Quality control: light scattering
G(t)=ádI(t)dI(t+t)ñ / áI(t)ñ2
[Link]
Quality control: SEC-MALS
The signal from the light-scattering detector is directly proportional to the molar mass of the
protein times the concentration (mg/ml). Combination of this signal with that from a
concentration detector (refractive index or absorbance) makes it possible to measure the
molar mass of each peak coming off the column.
Unlike conventional SEC methods, these molar masses from light scattering are independent
of the elution volume. Thus this technique can be used with "sticky" proteins that elute
unusually late as a result of their interactions with the column matrix, and also with highly
elongated proteins which elute unusually early for their molar mass. The molar masses derived
by this technique are generally accurate to 3% or better.
[Link]
[Link]
Quality control: analytical ultracentrifugation
Measure of the particle (biomolecules) distribution in presence of a centrifugal force.
Centrifugal force: F=mw2r
w=60000 rpm; r=6-7cm =>300 000g
[Link]
Quality control: analytical ultracentrifugation
Time t0 =0 Time = t
Quality control: analytical ultracentrifugation
]
D
15min.
O
Sedimentation velocity
[
e
30min.
c
n
45min.
0.8
a
0.7
b
Absorbance
r
0.6
o
s 0.5
b
0.4
a
0.3
0.2
0.1
-0.1
6,0 6,4 6,8
6.2 6.3 6.4 6.5 6.6 6.7 6.8 6.9 7
7,0
t: 1h radius (cm)
Centrifuge speed: High, molecules will sediment completely
Duration: several hours
Analysis: Front of sedimentation as a function of time
Sample: 400 µl or 100 µl (need for measurable UV absorption)
==>> Mass and shape (mass or conformational heterogeneity, aggregation,
oligomerization state / equilibria …)
Quality control: analytical ultracentrifugation
Sedimentation equilibrium
Absorbance
t: 24h
t: 0h
t: 0h
5,8 5,9 6,0 6,1
t: 24h Distance to the rotor axis (cm)
Centrifuge speed: Lower, molecules will not form a sedimentation front
Analysis: Equilibrium conditions resulting in a concentration gradient
Duration: Several days
Sample: 100 µl or 25 µl
Attention!!! for the analysis, UV absorption has to be proportional to the protein's
concentration
==>> Molecular mass with a precision of 1-2 %, indepent of the molecular
shape, chemical equilibrium -> KD (10-9-10-3 M), stoechiometry of molecular
complexes …
Quality control / initial protein characterization
7. Thermal Shift Assay (TSA)
checking relative stability of protein structure
can be used for optimization of experimental conditions (buffer composition),
addition of stabilizing small molecules etc
Quality control
8. Nuclear magnetic resonance (NMR)
degree of structuration from 1D or 2D 1H/15N-HSQC NMR spectra
peak => NMR signal of a proton in a given chemical environment
Folded
A well folded protein:
Each nucleus (1H) is situated in a distinct, well-
defined environment -> different resonance
frequencies
An unstructured protein:
Unfolded
Each nucleus is ≈ in an identical environment ->
same resonance frequencies for all nuclei of the
same type ([Link]. alanine Ca)
A floppy protein:
NMR spectrum = ∑ of the signals from individual
molecules. If each molecule is slightly different, the
corresponding resoancen frequencies will slightly
differ and the sum will correspond to broad signals
(peaks).
Quality control: NMR
correctly folded mutant proteins
mutant unable to adopt the
correct fold
Quality control (NMR): folded or unfolded ?
Pressure dependence of the [1H-
15N] HSQC spectrum recorded at
600 MHz on D+PHS Staphylococcal
Nuclease at 298 K
from:
Roche et al. Progress in Nuclear Magnetic Resonance
Spectroscopy 102–103 (2017) 15–31