vix.ing · top · new · best · stats · spec

The First Sequence: Fred Sanger and Insulin

2002/10/01 by Antony O.W. Stretton · 3 citations
Biochemistry, Genetics and Molecular Biology · Medicine · #Genomics and Rare Diseases #Hemoglobinopathies and Related Disorders

paper · pdf · doi:10.1093/genetics/162.2.527

openalex publication_date 2002/10/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/29

Abstract

FRED Sanger is an amazingly modest man, and his own retrospective, written after he retired, a delightful prefatory chapter for the Annual Reviews of Biochemistry, is called “Sequences, sequences, and sequences” (Sanger 1988). In it he describes the paths that led to the successful methods he developed for the sequencing of proteins, then RNA, and then DNA. What a career! Especially now, with the human genome largely finished, it is almost impossible to imagine a world without sequences of proteins and of nucleic acids. The fact that it has been only 50 years since Sanger showed that there was such a thing as the unique amino acid sequence of a protein seems amazing from where we now stand— sequences are such a dominant part of the world we work in today. Amazing, but perhaps not surprising, when we think of, say, the change in the power of computers that has occurred over just about the same time period. (I would love to be around, marbles intact, 50 years from now, to find out how the brain really works.) Before Fred’s work, it was already known that different proteins had different amino acid compositions, different biological activities, and different physical properties and that genes had an important role in controlling them. But in a world of biochemistry dominated by the role of enzymes in intermediary metabolism, it was not at all clear how molecules as large as proteins could be synthesized; the idea that proteins were stochastic molecules, with a sort of “center of gravity” of structure but with appreciable microheterogeneity, was taken seriously. This is the paradigm that Fred’s results shifted. This essay is a celebration of his first triumph: the first complete amino acid sequence determination of a protein, the B chain of insulin, published just over 50 years ago. Although I was not involved in the work in any way, I was scientifically aware enough to feel its contemporary impact. In 1957, Vernon Ingram, who at the time was working in the MRC Unit for the Study of the Molecular Structure of Biological Systems in the Cavendish Laboratory at Cambridge (later this morphed into the Laboratory of Molecular Biology), took me on as a research student. Vernon had just shown that sickle-cell hemoglobin differed from normal hemoglobin by a single amino acid substitution, the first characterization of the molecular consequences of mutation on proteins (Ingram 1956, 1957); earlier Neel (1949) had shown that sickle-cell anemia is inherited as a Mendelian character, and Pauling et al. (1949) showed that sickle-cell hemoglobin differed electrophoretically from normal hemoglobin and coined the term “molecular disease.” I had been trained as a chemist and knew nothing about proteins: I had heard Alex Todd’s exciting lectures to chemistry students on the organic chemistry of natural products that included vitamins, steroids, and nucleic acids, but not a word about proteins! To work on hemoglobin in Vernon’s laboratory, I was going to have to learn to become a protein chemist, and Francis Crick urged me to go and see Fred Sanger, who was at that time in the Biochemistry Department. In preparation for this, I read Fred’s 1952 review in Advances in Protein Chemistry entitled “The arrangement of amino acids in proteins” (Sanger 1952). I had also read a preprint of Francis’s own review, “On protein synthesis” (Crick 1958). Each review was completely stunning, but in very different ways. Francis’s article was full of elegant theory, with great leaps of speculation (at the time the existence of messenger RNA and transfer RNA was not known), but conveying a sense of joy that was as thrilling as that emerging from the Origin of Species. Fred’s was full of technique, and his review was scholarly, with major emphasis on experiments. Both of them made a huge impression on me, but I was particularly struck by Fred’s emphasis on the importance of developing new techniques. Today, when “hypothesis-driven research” has somehow become a gold standard, in many areas one must be very brave to admit to working on techniques, at least in the biological sciences. I think it may be different in chemistry—at least that is what my recent contacts with mass spectrometry suggest. Bill Dove has noted that Al Hershey (Nobel Prize winner in 1969), another temperamentally modest scientist who made seminal contributions, was also passionate about methods: “There’s nothing like technical progress! Ideas come and go, but technical progress cannot be taken away” (Hershey, quoted in Dove 1987). Another profound lesson I learned was that completely different scientific styles can be equally successful and valid and that personality has little to do with success. Again, the contrast between Fred and Francis is vast. After as little as half a minute with Francis, you know he has an exceptional mind, but with Fred it is much more subtle. He was brought up as a Quaker, which probably has a lot to do with his low-key and quiet demeanor, and during World War II he was a conscientious objector. In his 1988 article he writes, “Of the three main activities involved in scientific research, thinking, talking, and doing, I much prefer the last and am probably best at it. I am all right at the thinking, but not much good at the talking” and “Unlike most of my scientific colleagues, I was not academically brilliant,” and this after he had won two Nobel Prizes! So much for academic brilliance. Another unusual aspect of Fred’s character is his ability to pick and nurture people. I first got to know Fred indirectly through his graduate student Mike Naughton. Vernon moved to MIT in 1958 and took me with him, and Mike joined Vernon’s lab as a post-doc during my second year there. Mike and I became good friends, and I learned a lot of Fred’s techniques from Mike, so in some way I feel like one of Fred’s scientific grandchildren. Mike is a big, gentle Irishman from the West, who loved to sing and tell awful jokes. He had been a schoolteacher, and after doing his National Service in the Royal Air Force, he joined Fred as a technician. Soon he was transformed into Fred’s Ph.D. student, working also with Brian Hartley on a beautiful piece of work showing that the sequences around the active sites of the pancreatic serine proteases are identical (Hartley et al. 1959). Fred has done the same thing for at least two other people that I know, each taken on as a technician and turned into a Ph.D.—Bart Barrell, who is one of the world’s best nucleic acids sequencers, and Alan Coulson, who later became John Sulston’s righthand man on the genome projects of both Worm and Human. I will quote extensively from the introduction to Fred’s 1952 review, because it sets the stage beautifully for what was happening, a reevaluation of the nature of proteins. Fred wrote: It has frequently been suggested that proteins may not be pure chemical entities but may consist of mixtures of closely related substances with no absolute unique structure. The chemical results obtained so far suggest that this is not the case, and that a protein is really a single chemical substance, each molecule of one protein being identical to every other molecule of the same pure protein. Another earlier model (Bergmann and Niemann 1938) had suggested that proteins had periodic arrangements of amino acids, but the sequence of insulin ruled that out too. He added, These results [the insulin sequence] would imply an absolute specificity for the mechanisms responsible for protein synthesis and this should be taken into account when considering such mechanisms. He concludes, It is certain that proteins are extremely complex molecules but they are no longer completely beyond the reach of the chemist, so that we may expect to see in the near future considerable advances in our knowledge of the chemistry of these substances which are the essence of living matter. Remember that this was written only one year after Fred and Hans Tuppy had solved the structure of the B chain of insulin. The accuracy of protein synthesis remained an issue for many years after that. The previous review of the covalent structure of proteins had been written by Synge in 1943 (Synge 1943). In 1952, Fred writes: ... up to that time [1943] only a few simple peptides had been clearly identified from proteins by the classical and rather laborious methods of organic chemistry and Synge concluded that “the main obstacle to progress in the study of protein structure by the methods of organic chemistry is inadequacy of technique!” Probably the greatest advance that has been made recently in this field was the development by Martin and Synge (1941) of the entirely new technique of partition chromatography. The great problem in peptide chemistry has always been to find methods of fractionating the extremely complex mixtures produced by the partial degradation of a protein. Older methods of fractional crystallization and precipitation with various reagents were as a rule inadequate to deal with these mixtures, and countercurrent methods of high resolving power, which could fractionate nonvolatile, water-soluble substances, were needed. Partition chromatography, especially in the form of paper chromatography (Consden et al. 1944), is such a method, so that it has already been possible to identify as breakdown products of proteins more peptides using this technique than had previously been identified by the classical methods of organic chemistry. During the last few years, work in this field has centered largely on the development of methods, so that this review will be more a consideration of techniques and their uses than a discussion of results, which are still rather few. N-terminal sequences of insulin: One of the main reasons Sanger chose insulin for this work is that it was one of the few proteins available in pure form, and it was available in gram quantities because of its medical importance. At the time, the physical chemical evidence suggested a molecular weight of about 12,000. Fred invented the N-terminal labeling method using 1:2:4 fluorodinitrobenzene (FDNB), which reacts with amino groups under mild conditions that avoid degradation of the polypeptide chain. After complete acid hydrolysis of the dinitrophenyl (DNP)-protein, the DNP groups remain attached to the N-terminal amino acid and can be isolated and identified. Fred showed that there were four N-terminal residues per 12K insulin molecule, two of which were glycine and two phenylalanine (Sanger 1945), suggesting that there were four polypeptide chains in the 12K molecule. Cysteine was present, so it was thought that the chains were held together by -S-S- bridges, and indeed after performic acid oxidation, which splits the -S-S- bridges, insulin could be fractionated by precipitation into an A fraction and a B fraction; the A fraction had N-terminal glycine, and the B fraction had phenylalanine (Sanger 1949). The two fractions had different amino acid compositions, and neither contained tryptophan. Later it became clear that the 12K molecule is a noncovalent dimer of the fundamental molecular unit (Harfenist and Craig 1952), comprising one A chain and one B chain (see Figure 1). —The structure of bovine insulin. The lack of tryptophan was particularly fortunate, because it degrades upon acid hydrolysis, and one of the most important methods Fred used to get at the structure of insulin, with great success, was partial acid hydrolysis, which splits the peptide bonds almost randomly (more about that later). In fact, the first amino acid sequences of insulin came from partial acid hydrolysis of DNP-labeled A and B fractions. Some DNP peptides can be extracted into ethyl acetate from acid solution and then separated by silica gel columns; since DNP compounds are usually yellow, this was real “chroma”tography. For the B fraction, these peptides turned out to be the DNP-labeled N-terminal Phe followed by one or more other amino acids. Fred identified the DNP-amino acid and the other amino acids in the peptide after complete acid hydrolysis and then assembled the N-terminal sequence Phe.Val.Asp.Glu. Among the peptides that contained Asp or Glu, several different peptides had the same amino acid composition, so he concluded that both Asp and Glu were amidated in the original sequence. Other DNP peptides were not extracted into ethyl acetate. They were peptides derived from internal sequences surrounding Lys, to which DNP was linked by the amino group on the side chain. Since they all had free N-terminal amino groups (liberated from internal peptide bonds by partial acid hydrolysis), they were positively charged in acid, which explains why they did not extract into organic solvents. Amino acid analysis and relabeling with FDNB to determine the end groups gave the internal sequence Thr.Pro.Lys.Ala. The A fraction yielded the N-terminal sequence Gly.Ileu.Val.Glu.Glu. This article (Sanger 1949) was pivotal—it showed for the first time that at least some of the amino acids were in a unique sequence in insulin. Furthermore, the A and B fractions each yielded a unique sequence, suggesting that there were only two, not four, species of peptide chain in insulin—an A chain that contained about 20 amino acids and a B chain with about 30 amino acids—and he already had the sequence of over a quarter of the B more this article showed that it should be possible in to determine the structure of each chain by the methods developed in this partial hydrolysis, of the end group and partial hydrolysis of the longer In it is the methods that are and the of the has to be to them. Sanger and Tuppy did many to a between the and the which on the later of the hydrolysis, where the of the and also their in the were The complete sequence of the B The B chain was first (Sanger and Tuppy acid hydrolysis of the chain yielded many more products to be and the problem of of pure peptides from such a complex was They used several To fractionate they used on they the so that the acid peptides the performic of could be separated from the and of the very of the After at peptides amino acids, they used in solution with several acids or to the on peptides were separated by that Synge had used These which contained between and were then to paper chromatography, and the peptides were by with they were to and by amino acid and analysis was since at this stage no and method for all the DNP-amino acids had been the exciting you can tell from the article that they loved the into a longer sequence. One was that the peptide bonds to or residues are particularly to acid hydrolysis and always so they no with or This the complete but they the sequence of of the B two and a a and an which included the The of the sequence on the of and (Sanger and Tuppy At the time this was very because it was that proteases could the synthesis as as the of peptide the on insulin that this was not work by and his in and using simple had shown that these enzymes at different at or and amino acids, and the peptides obtained from of insulin was but to around amino acids and but also at several other the of the is in the protein and the on the of insulin were the first real The main of is that and at few but do so so the of the to be separated is The protein is into that can be isolated and by the same methods as in Sanger and Tuppy in some of the peptides from one were with a second to identify in the different of This time they obtained enough sequences sequences and to the and the sequence of the 30 amino acids in the B chain was These techniques, especially the of sets of peptides derived from different became the method for protein sequence determination for many Fred’s to sequencing by was for proteins, because of the of of complex At the time this work was the degradation technique had already been This method amino acids from the it later was the method of for the sequencing of peptides and proteins and became and at the In his Fred that he did not it because the products were not like the and so were to in the of and fraction The DNP compounds could be as the The A years Sanger and the sequence of the A amino acids with 30 in the B was more Again, they used partial acid hydrolysis of the chain and then to In these a new method was in silica a method that was very for these but was and by the paper methods for the DNP amino acids were so now they could positively identify the of These partial acid could be assembled into longer an that contained the known N-terminal sequence, a a and a which together included all the amino acids in the A but the of the peptide to serine and the and produced the and the sequence was with by A They this It would that no can be from these results the which the arrangement of the residues in protein In fact, it would more that there are no such but that each protein has its own unique an arrangement which it with its properties and and it for the that it in This is the that is so and had so much on the of molecular The -S-S- The A chain and B chain by are Sanger and his and on to another very important piece of covalent the three bonds in the molecule. They had to new methods, because the bonds under some of the conditions they used to peptide They that reagents like the under the conditions used for with pancreatic One of the bonds was but the A chain two and they no that could the peptide between them. showed that under acid conditions was by so they were to with their original sequencing techniques by partial acid the of partial acid hydrolysis products was now more because the mixtures were more complex different products joined together by the now, paper had been to the methods, by using different conditions and the with paper chromatography in various they the that gave the of the bonds et al. were two and one in the A chain (see Figure 1). The Nobel was to the importance of this work, and Fred’s first Prize came in 1958 second in was for sequencing nucleic the methods used in this sequence determination were largely Amino acid analysis was done by the by with and with peptides that is good A little and and their developed methods for peptides on et al. and for amino acid a that was et al. 1958). own on human hemoglobin and protein, was I isolated peptides by paper and chromatography, but used amino acid analysis to determine the After a I have recently to peptide and it is a different is now done by high chromatography which has and sequencing is and at least three of more and in most it is done by using a big, in one or another lot of the is mass spectrometry is of importance in The peptides can be randomly in and the by molecular mass to the sequence. idea as partial acid One problem is that and have the same So do and but they are and it is to identify by the mass after The change is that most protein sequencing is done indirectly from sequences, and only of who work on proteins and peptides in my are still to work with the amino acids. on a very different of who did sequencing using the original methods would probably a really group to because we were to a of organic in the on both by and chromatography. For the paper was in which as a was with in in the and it was to avoid it on as you the was by a The most used contained and were really The to be separated was the paper in a and with a then the was to the which was out on a The was in the paper on each side so that the at the at the same time, and this a and was best not done right after or This the into a very you did it but also you for a The for paper chromatography were also and on the when I used acid, my would and I would have a the In we were of the of this sort of It is also to see the of the in the of that the covalent structure of insulin. to Fred’s development of the FDNB method of N-terminal labeling as and indeed it but his of chemical and methods and his of new methods for mixtures of were also for his success. chromatography was a new technique that he and his used but chromatography was not for the complete of the complex mixtures by partial acid hydrolysis, so they used various methods, to peptides amino acids, in solution or in silica and in the et al. they to paper a technique that separated peptides almost on the of and molecular It was more when used in with chromatography, which was also to technique was the of Vernon method for normal and sickle-cell Fred stunning, Hans on the sequence of the B chain of insulin a huge particularly The of was already and a few years later showed that sites of mutation a single also a Fred had another that of The molecular the two and the fact that there was a sequence in proteins led to the thought that there had to be a Fred came and the other two and in the of protein but that is another I am very to for and as is and

Citations

Cited by