★ wanayoo — archive 1999 http://evolution.genetics.washington.edu/phylip/software.etc1.htmlNouvelle recherche | Portail wanayoo
To go to top of Software page

To previous part of Software page


Jun Adachi and Masami Hasegawa have written a package MOLPHY 2.2, carrying out maximum likelihood inference of phylogenies for either nucleotide sequences or protein sequences. Their protein sequence maximum likelihood program, ProtML, is a successor to the one they made available to me for distribution on a nonsupported basis in PHYLIP, and is much improved over that. It is one of two protein maximum likelihood programs available. The package is distributed free in C source code, with documentation, by ftp from sunmh.ism.ac.jp. An executable version for Windows95 or Windows NT on Intel processors, and also one that works on Windows NT on DEC Alpha processors, is available from Russell Malmberg at the Botany Department of the University of Georgia (russell@dogwood.botany.uga.edu) by World Wide Web at http://dogwood.botany.uga.edu/malmberg/software.html


Gary Olsen, of the Department of Microbiology, University of Illinois, Urbana, Illinois (gary@phylo.life.uiuc.edu) has developed a speeded-up replacement for my program DNAML coded in C, called fastDNAml. It achieves a number of economies and also is organized so that it can be run on parallel processors -- he and his co-workers have constructed trees of very large size on a high-speed parallel processor. The program can be compiled using the "p4" portable parallel processing toolkit. It can also be run in ordinary serial mode on workstations where it is faster than DNAML.


Denis Beaumont (beaumont@transpac.atlas.fr) has made a parallelized version of fastDNAml called VeryfastDNAml. It is parallelized with the TreadMarks distributed shared memory system, which is a not-quite-free environment for parallelization that runs on many workstation-class machines. The C source code of VeryfastDNAml is available by ftp from the Institut Pasteur server ftp.pasteur.fr in directory /pub/GenSoft/unix/evolution/FastDNAml as file fastDNAml-tmk.tar.gz. There is a web page access to this ftp distribution at http://bioweb.pasteur.fr/seqanal/soft-pasteur.html#veryfastdnaml, which includes a link to the TreadMarks project.


Ziheng Yang of the Department of Genetics and Biometry, University College London, (z.yang@ucl.ac.uk) has released PAML, version 2.0g, a package of programs for the maximum likelihood analysis of nucleotide or protein sequences, including codon-based methods that take into account both amino acids and nucleotides. The programs can estimate branch lengths in a phylogenetic tree and parameters in the evolutionary model such as the transition/transversion rate ratio, the gamma parameter for variable substitution rates among sites, rate parameters for different genes, and synonymous and nonsynonymous substitution rates. They can also test evolutionary models, calculate substitution rates at particular sites, reconstruct ancestral nucleotide or amino acid sequences, simulate DNA and protein sequence evolution, compute distances based on the synonymous and nonsynonymous changes, and of course do phylogenetic tree reconstruction by maximum likelihood and Bayesian Markov Chain Monte Carlo methods. The strength of the package lies in its rich implementation of evolutionary models, though Yang coments that tree-making is not a strong point of the current version. The autocorrelated gamma distribution implementation is analogous to the Hidden Markov Model scheme available in PHYLIP. The package is available as ANSI C source code for Unix systems, as PowerMac executables and as executables that run on Windows95, Windows98, and WindowsNT. See the PAML web page at http://abacus.gene.ucl.ac.uk/ziheng/paml.html. It can be downloaded by ftp from abacus.gene.ucl.ac.uk in directory pub/paml.


Bret Larget and Donald Simon of the Department of Mathematics and Computer Science, Duquesne University, Pittsburgh, Pennsylvania (bambe@mathcs.duq.edu) have written BAMBE (Bayesian Analysis in Molecular Biology and Evolution) version 2.01 beta, a program for Bayesian analysis of phylogenies with DNA sequence data. It uses a prior distribution of trees and arearrangement mechanism introduced in the paper: Mau, B., M. A. Newton, and B. Larget. 1997. Bayesian phylogenetic inference via Markov chain Monte Carlo methods. Molecular Biology and Evolution 14: 717-724. The trees and parameter values are sampled by a Metropolis algorithm Markov Chain Monte Carlo sampling. The resulting posterior distribution can be used to characterize the uncertainty about not only the tree, but the parameters of the substitution model as well. The program is in C++ source code for Unix, and is distributed from its web page at http://www.mathcs.duq.edu/larget/bambe.html.


[Modeltest icon] David Posada and Keith Crandall at the Department of Biology, Brigham Young University (dp47@email.byu.edu) has released Modeltest version 3.0, a program to test a hierarchy of statistical models of DNA evolution using the Likelihood Ratio Test criterion and the AIC (Akaike Information Criterion). The likelihood values are obtained by running PAUP*. MODELTEST accepts likelihood scores corresponding to 56 models of DNA substitution including whether transition and transversion rates are equal, whether rates at different sites are equal, and whether there are invariant sites. Modeltest is described in the paper: Posada, D. and K. A. Crandall. 1998. MODELTEST: testing the model of DNA substitution. Bioinformatics 14: 817-818. It is available as executables for Macintosh and PowerMac, for Windows95/98/NT, for Linux and for Suns, and in source code as C for the Metrowerks C compiler. It is distributed from its web site at http://bioag.byu.edu/zoology/crandall_lab/modeltest.htm.


[PLATO icon here] Nick Grassly, currently of the Zoologisches Institut, Universität München (grassly@zi.biologie.uni-muenchen.de), has written PLATO, version 2.11, a program that takes sequential PHYLIP-style DNA sequences followed by their maximum likelihood phylogeny, and using a likelihood approach with sliding window analysis and Monte Carlo simulation of the null distribution detects anomalously evolving regions in the DNA sequences and assesses their significance. This may lead to the detection of, for example, recombination, gene conversion or convergence, or reveal variable selective pressures along the gene sequence. A general substitution model is used that can allow the test to reveal differences due to recombination while ignoring those due to varying rate of evolution. The method is described in the paper: Grassly, N. C., and E. C. Holmes. 1997. A likelihood method for the detection of selection and recombination using sequence data. Molecular Biology and Evolution 14: 239-247. It is available for Macintoshes (including PowerMacs) or in source code for Unix systems. It requires substantial amounts off memory, especially when sequences analysed are long. Use of Power Macintoshes or UNIX systems is recommended. It is distributed free from the University of Oxford Zoology Web server at http://evolve.zoo.ox.ac.uk/Plato/Plato2.html.


Mika Salminen and Wayne Cobb (msalminen@hiv.hjf.org and wcobb@reed.hjf.org), of the Henry M. Jackson Foundations for the Advancement of Military Medicine, Walter Reed Army Institite of Research, Bethesda, Maryland, have released the Bootscanning Package, version 1.0beta. This is a series of shell scripts and programs that analyze DNA sequences for evidence of recombination. It breaks the sequence into separate pieces that are analyzed for the bootstrap support of various groups, and it looks for evidence of significant conflict among trees for different parts of the sequence. The programs are currently available only as Sun executables. They require GDE 2.2a and PHYLIP version 3.4 to work. They are available by anonymous ftp from from http://www.ktl.fi in directory /hiv/mirrors/pub/programs.


Gráinne McGuire and Frank Wright (grainne@bioss.sari.ac.uk and frank@bioss.sari.ac.uk) of Biomathematics and Statistics Scotland, in Dundee, have released TOPAL, which checks for evidence of past recombination events, by looking for changes in the inferred phylogenetic tree TOPology between adjacent regions of a multiple sequence ALignment. Their method detects recombinations by sliding a window along a sequence alignment, and measuring the discrepancy between the trees suggested by the first and second halves of the window, using distance matrix methods. This is described in the paper: McGuire, G., F. Wright, and M. J. Prentice. 1997. A graphical method for detecting recombination in phylogenetic data sets. Molecular Biology and Evolution 14: 1125-1131. The TOPAL program is also described in a paper: McGuire, G. and F. Wright. 1997. TOPAL: recombination detection in DNA and protein sequences. Bioinformatics 14: 219-220. TOPAL is a set of Unix Bourne shell scripts and C code, plus four programs in C from my PHYLIP package. These are available from the TOPAL web site at http://www.bioss.sari.ac.uk/~frank/Genetics/topal.html.


Ingrid Jakobsen and Simon Easteal of Australian National University, Canberra, have released reticulate. (Ingrid Jakobsen is now at the Institute of Molecular Evolutionary Genetics, Pennsylvania State University, and her e-mail address is ibj1@psu.edu). It is a compatibility matrix program for DNA sequences that has features designed to test for evidence of reticulate evolution (such as recombination). The program computes and displays a pairwise compatibility matrix for all pairs of sites. It can randomize the order of sites and compute the fraction of compatible sites in a region for the randomizations, to test whether there is a pattern suggesting reticulation. The program is distributed as C source code for Unix and X Windows, though there are some limited ways of running it without X Windows. It is described in the paper: Jakobsen, I. B. and S. Easteal. 1996. A program for calculating and displaying compatibility matrices as an aid in determining reticulate evolution in molecular sequences. CABIOS 12: 291-295. It is available from its web site at http://jcsmr.anu.edu.au/dmm/humgen/ingrid/reticulate.htm.


Kim Fisker (kfisker@daimi.aau.dk) of the Computer Science Department at Aarhus University, Denmark has released RecPars, which does a parsimony analysis of DNA sequences. It tries to find the best phylogenies for different regions of the sequences and thereby postulating a recombination event between these segments. The method is described in a paper: Hein, J. 1993. A heuristic method to reconstruct the history of sequences subject to recombination. Journal of Molecular Evolution 36: 396-406. RecPars is available as C source code for Unix. It is distributed by ftp from ftp.daimi.aau.dk in directory pub/empl/kfisker/programs/RecPars.


John Maynard Smith and Noel Smith of the School of Biological Sciences of the University of Sussex (noelsmith@yahoo.com) have released programs to carry out their homoplasy test for recombination in sequences. The test is described in a paper: Maynard Smith, J. and N. H. Smith. 1998. Detecting recombination from gene trees. Molecular Biology and Evolution 15: 590-599. The programs are distributed in QBASIC for DOS and must be run using QBASIC. They are available from Maynard Smith's web site at http://www.biols.susx.ac.uk/Home/John_Maynard_Smith/.


[LARD icon here] Andrew Rambaut of the Department of Zoology, University of Oxford, England (andrew.rambaut@zoo.ox.ac.uk) has produced LARD (Likelihood Analysis of Recombination in DNA) version 2.2, a program to detect the presence of recombination in a set of sequences. LARD looks at the set of sequences to discover which are the most plausible parents of a potentially recombinant sequence, and performs a likelihood ratio test for each possible breakpoint position of whether the three-species tree differs on the two sides of the breakpoint. LARD is described as an extension of a method suggested by John Maynard Smith: Maynard Smith, J. 1992. Analysing the mosaic structure of genes. Journal of Molecular Evolution 34: 126-129. It is described in a paper: Holmes, E. C., M. Worobey, and A. Rambaut. 1999. Phylogenetic evidence for recombination in dengue virus. Molecular Biology and Evolution 16: 405-409. LARD is available as C source code and as a Macintosh executable from its web site at http://evolve.zoo.ox.ac.uk/Lard/Lard.html.


[SPOT icon here] Andrew Rambaut of the Department of Zoology, University of Oxford, (andrew.rambaut@zoo.ox.ac.uk) and Nick Grassly, currently of the Zoologisches Institut, Universität München (grassly@zi.biologie.uni-muenchen.de), have written SPOT (Sequence Parameters Of Trees). SPOT is a program that will calculate the likelihood of a given tree topology for a set of aligned nucleotide sequences. For each topology, SPOT will estimate the maximum likelihood values of branch lengths and other parameters of the model of nucleotide evolution that has been chosen. Such parameters include the ratio of transitions to transversions (TS/TV ratio) and relative rates of substitution at different codon positions. Branch lengths can also be constrained to assume a molecular clock hypothesis. Multiple datasets and multiple trees can be analysed which is useful for performing Monte Carlo simulations of hypothesis (parametric bootstraps). Although SPOT does not estimate tree topology, an accompanying program, SPOTSHELL, will iterate between fastDNAml and SPOT until the maximum likelihood parameters and topology has been found (or at least something close to it). SPOT is available as C source code for Unix workstations, or as Macintosh sources and executables. It can be obtained from the SPOT Web page at http://evolve.zoo.ox.ac.uk/Spot/Spot.html.


Gary Olsen of the Department of Microbiology, University of Illinois, Urbana, Illinois (gary@phylo.life.uiuc.edu) has written dnarates version 1.0. It reads a set of DNA sequences and a tree, and for that tree makes a maximum likelihood estimate of the rate of evolution at each site. This is done by taking the rate at each site as a separate parameter and maximizing the likelihood with respect to all those parameters. The program is available as generic C source code. It is based in part (with my permission) on code from my PHYLIP program DNAML. dnarates is available by ftp from the IUBIO ftp server at ftp://rdp.life.uiuc.edu/pub/RDP/programs/DNArates/.


[Spectrum icon here] Mike Charleston (mcharles@udcf.gla.ac.uk) of the Division of Environmental and Evolutionary Biology of the University of Glasgow has developed Spectrum, a program for finding bipartition spectra from phylogenetic molecular and distance data, according to the method of Hendy et al. (1994) (Hadamard transforms) for moderately sized data sets (up to 18 taxa). The program also implements a branch-and-bound search for the "closest tree" - that is, the tree whose expected spectrum is closest to the spectrum derived from the observed PowerMac, 68k Macintosh, and Windows95 or Windows NT executables are available from its Web site in the Glasgow Taxonomy web pages: http://taxonomy.zoology.gla.ac.uk/~mac/spectrum/spectrum.html.


Ingrid Jakobsen, Susan Wilson, and Simon Easteal, of Australian National University, Canberra, have released partimatrix. (Ingrid Jakobsen is now at the Institute of Molecular Evolutionary Genetics, Pennsylvania State University, and her e-mail address is ibj1@psu.edu). This program computes a "partition matrix" from aligned DNA sequence data. The method finds partitions of the sequences into two groups and presents a matrix which describes the conflict and agreement among these partitions. The objective is to discover parts of the DNA sequence which imply different trees. It is described in the paper by I. B. Jakobsen, S. R. Wilson and S. Easteal. 1997. The Partition Matrix: Exploring variable phylogenetic signals along nucleotide sequence alignments. Molecular Biology and Evolution 14: 474-484. The program is distributed as C source code for Unix systems with X Windows. It is available from its web site at http://jcsmr.anu.edu.au/dmm/humgen/ingrid/partimatrix.htm .


Pablo Goloboff, of INSUE - Fundación e Instituto Miguel Lillo 205, 4000 S. M. de Tucumán, Argentina, has written Nona (Noname), PiWe (Parsimony with Implied WEights), and SPA to carry out parsimony including weighted parsimony analyses. Nona searches for most parsimonious trees according to character weights defined by the user a priori. Pee-Wee calculates weights of the characters by a method introduced by Goloboff, a noniterative version of J. S. Farris's "successive weighting". It was described in Goloboff's paper in Cladistics 9: 83-91, 1993. SPA is a generalized parsimony program that allows differential weighting of changes between different states. Nona is said to be faster than other parsimony programs. A Windows 95/98/NT version of Nona which includes the functionality of Piwe and SPA is available as shareware (with a free 30-day trial period) from its web page at http://www.cladistics.com/about_nona.htm. The shareware fee of $40 should be paid to the author at the above address or to James M. Carpenter, Department of Entomology, American Museum of Natural History, Central Park West at 79th Street, New York, NY 10024. Send the money and the name in which the copies are to be registered. An earlier demo version of these programs which runs on DOS is also available. It requires the user to hit an extra key each time they execute a command. It is can be fetched by anonymous from ftp.vims.edu in directory other/hennig as file pars-pag.exe, a self-extracting DOS Zip file.


Yasuo Ina of the National Institute of Agrobiological Resources, Tsukuba, Japan ha developed ODEN, a package of programs for doing distance matrix analyses on nucleotide or protein sequences. It is described in CABIOS 10: 11-12 (1994). It is available free by anonymous ftp from directory pub/unix/oden on ftp.dna.affrc.go.jp as C source code for Unix systems.


A. Luettke and R. Fuchs have written MacT, a package of programs for Macintoshes that compute distances and compute Neighbor-Joining phylogenies for them. The programs work on 4 through 26 sequences, and source code in Microsoft QuickBasic is provided as well as compiled executables. The package is free and is available on the molecular biology software servers. For example, it is available on by anonymous ftp on the Indiana University IUBIO server ftp.bio.indiana.edu it will be found in directory molbio/mac. The programs are described in CABIOS 8: 591-594, 1992.


Andrey A. Zharkikh, Andrey Rzhetsky, and co-workers in the Institute of Cytology and Genetics, Siberian Branch of the Russian Academy of Sciences, Novosibirsk, Russia, have produced VOSTORG, a package of programs for alignment (both manual and automatic) and inferring phylogenies by distance methods and parsimony for molecular sequences. (Zharkikh and Rzhetsky are currently in the US; their e-mail addresses are zharkikh@myriad.com and andrey@genome2.cpmc.columbia.edu). VOSTORG runs on under DOS on PC-compatibles and includes some rather fancy graphics (for DOS). It is available from its Web page in Russia from http://molevol.bionet.nsc.ru/vs.htm. The programs are described in a paper: Zharkikh, A. A., A-Yu. Rzhetsky, P. S. Morosov, T. L. Sitnikova, and J. S. Krushkal. 1991. VOSTORG: a package of microcomputer programs for sequence analysis and construction of phylogenetic trees. Gene 101: 251-254.


Walter Fitch (wfitch@uci.edu), of the Department of Ecology and Evolutionary Biology, of the University of California at Irvine, has available by anonymous ftp at daedalus.bio.uci.edu in directory pub/outgoing/evoprog about 20 programs which carry out various kinds of phylogeny estimation and related tasks. They are available in source code in FORTRAN 77, (except for a few which are in C) and also as Sun SPARC executables and as DOS executables. They include:

There are also many programs that convert sequences among various formats, generate all possible trees, shuffle sequences, align sequences, and do various other functions. The programs are available by anonymous ftp from daedalus.bio.uci.edu in directory pub/outgoing/evoprog.

There is also TDRAW which draws a tree in Postscript. This program is in C, and is not available as a DOS executable. It is available in directory pub/outgoing/tdraw.


Nicholas Galtier of the University of Lyon (galtier@biomserv.univ-lyon1.fr) has written Phylo_win, a "graphic interface" for molecular phylogenetic inference. It performs neighbor-joining, parsimony and maximum likelihood methods and can bootstrap with any of them. Many distances can be used including Jukes & Cantor, Kimura, Tajima & Nei, Galtier & Gouy (1995), LogDet for nucleotidic sequences, Poisson correction for protein sequences, Ka and Ks for codon sequences. Species and sites to include in the analysis are selected by mouse. Reconstructed trees can be drawn, edited, printed, stored, evaluated according to numerous criteria. Taxonomic species groups and sets of conserved regions can be defined by mouse in both tools and stored into sequence files, thus avoiding multiple data files. It is entirely mouse-driven. Most usual sequence file formats are read: CLUSTAL, FASTA, PHYLIP, MASE. It runs under X windows on many Unix workstations. It is described in the paper: Galtier, N., M. Gouy, and C. Gautier. 1996. SeaView and Phylo_win, two graphic tools for sequence alignment and molecular phylogeny. Computer Applications in the Biosciences 12: 543-548. It is distributed as C source code (to compile it one needs the NCBI Vibrant tool kit). It is also available as executables for SunOS, Solaris, SGI Unix, IBM RISC Unix, Linux, HP/UX, and DEC Alpha (Digital Unix). It can be fetched from its web page at http://pbil.univ-lyon1.fr/software/phylowin.html. It can also be obtained by anonymous ftp from biom3.univ-lyon1.fr in directory pub/mol_phylogeny. A PC Linux executable is available at http://evolution.bmc.uu.se/~thomas/mol_linux. A Digital OpenVMS executable is also available as http://seqaxp.bio.caltech.edu:8000/pub/SOFTWARE/phylo_win_vms.zip.


F. James Rohlf has written NTSYSpc (Numerical Taxonomy System, Version 2.0), a clustering program that includes calculation of various kinds of distance measures, as well as Hierarchical clustering methods such as UPGMA as well as Neighbor-Joining and consensus trees. It can also do a variety of other things including ordination, scatter diagrams, and elliptic Fourier transforms (for shape analysis). NTSYSpc 2.0 is a Windows95 executable which will also run on Windows NT. It is available for $275 ($210 for educational and government institutions). 10-user site licensese are also available. It is distrubuted by Exeter Software (the biological software company, not the warehouse-inventory-software house of the same name). Their e-mail address is sales@exetersoftware.com. Their toll-free telephone number is 800-842-5892, their not-so-free phone number is +1-631-689-7838, and their fax number is +1-631-689-0103. Their mailing address is 100 North Country Road, Setauket, NY 11733-1345 USA. FAX, or phone (toll-free telephone within the USA). Further information is available on their Web page at http://www.exetersoftware.com/cat/ntsyspc.html.


[CAFCA icon here] Rino Zandee (zandee@rulsfb.leidenuniv.nl), of the Institute of Evolutionary and Ecological Science, Van der Klaauw Laboratory, Leiden University, has written CAFCA version 1.5j, the Collection of APL Functions for Comparative Analysis. It carries out a search for the most parsimonious tree with discrete-character data (either two-state or multistate), using a search for cliques of component compatibility (monothetic subsets) to propose the candidates for most parsimonious trees. The program is written as functions in the APL language, but Macintosh and PowerMac executables are distributed. The program is free and is available from the CAFCA Web Site http://wwwbio.leidenuniv.nl/~zandee/cafca.html.


[TREE-PUZZLE icon here] Korbinian Strimmer(http://users.ox.ac.uk/~strimmer) now at the Department of Zoology, University of Oxford, U.K.), and Arndt von Haeseler (haeseler@eva.mpg.de) now at the Max-Planck-Institute for Evolutionary Anthropology, Leipzig, (both previously of the Zoologisches Institut of the Universität München) have developed TREE-PUZZLE version 4.0.2, (formerly called PUZZLE) a program for maximum likelihood analysis for nucleotide and amino acid alignments. It infers phylogenies by "quartet puzzling", a method that applies maximum likelihood tree reconstruction to all possible quartets of taxa and subsequently tries to combine most of the four-taxa maximum likelihood trees to construct an overall maximum likelihood tree. Usually there are several possible solutions. A consensus tree generated from the quartet puzzling trees shows nodes that are well supported. More details about the algorithm and on the phylogenetic accuracy can be found in the papers: K. Strimmer and A. von Haeseler. 1996. Molecular Biology and Evolution 13: 964-969 and K. Strimmer, N. Goldman, and A. von Haeseler. 1997. Molecular Biology and Evolution 14: 210-211. TREE-PUZZLE supports all popular models of sequence evolution of nucleotides and proteins, and can take rate heterogeneity among sites into account. It computes pairwise maximum likelihood distances for many different models of sequence evolution (TN, HKY, F84, SH, Dayhoff, JTT, mtREV24 and BLOSUM62), and estimates parameters of the models. It can estimate maximum-likelihood branch-lengths for user-specified trees and perform likelihood ratio tests of clockness as well as Kishino-Hasegawa-Templeton tests. The program is written in ANSI C and is compatible with PHYLIP files. precompiled executables are distributed for PowerMac and for Windows 95/98/NT. For UNIX and VMS systems files for automated compilation are provided. It is available from the TREE-PUZZLE web page at http://www.tree-puzzle.de or by anonymous ftp from:

Its online manual can be viewed at http://www.tree-puzzle.de/manual.html. A Debian Linux package of TREE-PUZZLE is available at its web site at http://www.debian.org/Packages/unstable/misc/puzzle.html.


Mike Holder (holder@mbl.edu) and Andrew Roger(roger@mbl.edu) of the Marine Biological Laboratory in Woods Hole, Massachusetts are distributing a shell script program for Unix systems, puzzleboot that allows the analysis of multiple data sets with TREE-PUZZLE. It is designed for use with the distance matrix option of TREE-PUZZLE, to make use of the distance calculation methods. It is available from the TREE-PUZZLE web page at http://www.tree-puzzle.de.


Kay Nieselt-Struwe (kns@phy.auckland.ac.nz) of the Department of Physics of the University of Auckland, New Zealand has released version 1.0 of STATGEOM. It carries out computation of the statistical geometry in distance and in sequence space of a set of aligned DNA/RNA, amino acid or binary sequences. The user can decide to either compute the overall tree-likeness of the whole set, or a certain subset, or given a tree of the sequences compute the reliability of certain edges in the tree. Postscript files of the graphs of the statistical geometry are automatically generated. A sequence reformatting utility allow various sequence formats to be read in. STATGEOM is written in ANSI C; source code with documentation and a Sun SPARC executable are available by anonymous ftp at cage.mpibpc.gwdg.de (or 134.76.209.64) in directory pub/kniesel. The method of statistical geometry was originally published in: Eigen, M., Winkler-Oswatitsch, R. and Dress, A. 1988. Statistical geometry in sequence space: a method of comparative sequence analysis. Proc. Natl. Acad. Sci. USA 85: 5913-5917.


[SplitsTree icon here] Rainer Wetzel and Daniel Huson of the University of Bielefeld (huson@mathematik.uni-bielefeld.de) have developed a program SplitsTree2 for SplitsTree2 carries out the split decomposition method of A. Bandelt and A. Dress, Bandelt and Dress's p-splits, and spectral analysis (Hendy, Penny, Szekely, Steel & Erdos). It can process sequence or restriction site data, and can do does bootstrapping. It also contains an implementation of Cooper, Penny and Steel's method for dating divergences and also their molecular clock test. (Molecular Phylogenetics 1: 242-252 (1992)). It is also discussed in the paper: Huson, D. H. 1998. SplitsTree: analyzing and visualizing evolutionary data. Bioinformatics 14: 68-73. SplitsTree2 is currently available as a Mac program or in a Unix version for a number of different machines (Sun, SGI, DEC and HP). These use the program ghostview to draw the computed graph. The Mac version draws the graph in its own window, and the picture can be copied and pasted or printed in the usual way. There is no Windows version at present. It is available by ftp from ftp.uni-bielefeld.de in directory pub/math/splits. SplitsTree2 is software under development. A server is also maintained which uses SplitsTree 2 to analyze data submitted via its web page.


Igor Kuznetov and Pavel Morozov (kuznets@mailhost.bionet.nsk.su) of the Institute of Cytology and Genetics, Novosibirsk, Russia, have produced GEOMETRY, a package for nucleotide sequence analysis using the method of statistical geometry in sequence space (M. Eigen, R. Winkler-Oswatitsch, and A. Dress. 1988. Statistical geometry in sequence space: A method of quantitative comparative sequence analysis, Proc. Natl. Acad. Sci. USA 85: 5913-5917). The program is described in an article by Kuznetsov and Morozov in 1996 in CABIOS 12: 297-301. The package uses the same data formats for sequence and tree input as the ones used in VOSTORG package. GEOMETRY is available as a DOS executable. It is available for downloading by Web from http://molevol.bionet.nsc.ru/soft.htm or by ftp from ftp.bionet.nsk.su in directory incoming/molevol and also from the EMBL file server ftp.ebi.ac.uk in directory pub/software/dos.


Vincent Berry of the Université Jean Monnet in St.-Etienne, France (vberry@univ-st-etienne.fr) has released PhyloQuart version 1.3, a package of programs inferring phylogenies from quartets. It is able to use either nucleotide sequences or distances. It implements the Q* method of tree reconstruction, which is inspired by the work of Bandelt and Dress, and is described in the forthcoming paper: Berry, V. and O. Gascuel. Inferring Evolutionary Trees with Strong Combinatorial Evidence. Theoretical Computer Science, to appear. PhyloQuart is available as C source code which can be compiled on Unix systems, from its web site at http://www.univ-st-etienne.fr/eurise/LOGICIELS/PHYLOQUART/main-pq.html. PhyloQuart is also available as a Web server from the server of the Institut Pasteur.


Stephen J. Willson (swillson@iastate.edu) of the Department of Mathematics, Iowa State University, has produced a package of programs to infer phylogenies from quartets of species. They infer phylogenies of individual quartets by parsimony, and in combining them use information on how strongly the phylogeny for that quartet is preferred over its alternatives, or by measures of how well the group fits into a given placement on a tree, as judged by quartets. The methods are described in two papers: Willson, S. J. 1998. Measuring inconsistency in phylogenetic trees, Journal of Theoretical Biology 190: 15-36, and Willson, S. J. 1998. Building phylogenetic trees from quartets by using local inconsistency measures . Molecular Biology and Evolution 16: 685-693. The programs are in C and are described as having successfully been compiled on PowerMac systems using the Codewarrior C compiler. PowerMac executables are also provided. The programs are available at Willson's software web site at http://www.public.iastate.edu/~swillson/software.html.


James Lake of the Department of Molecular, Cell and Developmental Biology of the University of California, Los Angeles (lake@mbi.ucla.edu) has released Gambit, which implements a method called Boostrapper's Gambit. The method involves bootstrap sampling sequences, computing trees for quartets of species, and assembling larger trees out of quartets that have significant boostrap support. One of the methods available to estimate trees from the quartets is paralinear (LogDet) distances. Other distance methods and parsimony are also available. The program is available as a DOS executable, free to noncommercial users on a trial basis until January 15, 2001. Commercial users are asked to pay $50 on a shareware basis. The program is available at its web site at http://www.lifesci.ucla.edu/mcdbio/Faculty/Lake/Research/Programs/.


Arne Röhl, Peter Forster, and Hans-Jürgen Bandelt (Forster is at pf223@cus.cam.ac.uk) have written Network 2.0b, a program to infer networks (which have more connections than trees). The networks are median-joining networks, a method which is described in a paper: Bandelt, H-J., P. Forster, and A. Röhl. 1999. Median-joining networks for inferring intraspecific phylogenies. Molecular Biology and Evolution 15: 1108-1114. The program is available as shareware (free until 1 June 2000) as a DOS executable from Fluxus Engineering at its web site at http://www.fluxus-engineering.com/sharenet.htm.


Pierre Rioux and Tim Littlejohn of the Informatics Division of the Organelle Genome Megasequencing Program at the Universite de Montreal has made Available PARBOOT, a program that takes bootstrap sampled data sets and splits them up, submitting each to a different computer, so as to run bootstrapping quickly on networks of computers. It is available free as C source code by ftp from megasun.bch.umontreal.ca in directory pub/parboot. It requires a networked system of computers with PHYLIP, a Perl interpreter, and appropriate accounts and permissions.


Naoko Takezaki (ntakezak@lab.nig.ac.jp) of the Center for Information Biology of the National Institute of Genetics, Mishima, Japan, has written Lintre (Phylogenetic tests of the molecular clock and linearized tree), a package of programs for Sun workstations. The programs include:

The two-cluster test is essentially the relative rate test for many sequences. The branch length test is the test of rate difference for each sequence under the tree root from the average rate of all sequences. The tests are described in: Takezaki, N., A. Rzhetsky, and M. Nei. 1995. Phylogenetic test of the molecular clock and linearized trees. Molecular Biology and Evolution 12: 823-33. The programs are available as C source code and also as DOS executables. The are distributed (as a compressed tar archive of the source code with examples and documentation, and also as a self-extracting archive of sources and DOS executable) from her web site at http://cib.nig.ac.jp/dda/ntakezak.html and by anonymous ftp from the IUBio server ftp.bio.indiana.edu in directory molbio/evolve.


Naoko Takezaki (ntakezak@lab.nig.ac.jp) of the Center for Information Biology, National Institute of Genetics, Mishima, Japan, has written njbafd which constructs a neighbor-joining tree or a UPGMA tree from microsatellite data and other allele frequency data. Bootstrapping can be carried out. The program includes Goldstein et al.'s distance for microsatellite loci. There are is a source code Unix version and an executable DOS version. They are available from her web site at http://cib.nig.ac.jp/dda/ntakezak.html. They are also available by ftp in directory molbio/evolve of the IUBIO archive.


[WET icon] Joaquin Dopazo of the Bioinformatics department of GlaxoWellcome SA, Spain (jd19662@glaxowellcome.co.uk) has written WET (Windows Easy Tree), version 1.3, which is an easy-to-use program for inferring phylogenies from sequence data by distance matrix methods. The main goal in the development of WET was to make a really user friendly program able to interact with other phylogenetic packages. WET can import files of a number of different formats. It calculates distances by a number of different methods and constructs phylogenetic trees using neighbor-joining, UPGMA and WPGMA procedures. It is a Windows95/98/NT executable. It is available from its web site at http://www.cnb.uam.es/~dopazo/software/wet.html.


David Penny (Institute of Molecular Biosciences, Massey University, Palmerston North, New Zealand) has been offering for free distribution two DOS programs, one a fast parsimony program, TurboTree. There is also another, Great Deluge, an approximate search for the most parsimonious tree by a quasi-random method. He tells me that funding exigiencies are such that he may soon have to start charging for these. His electronic mail address is dpenny@massey.ac.nz.


David Penny of the Institute of Molecular Biosciences, Massey University, Palmerston North, New Zealand (dpenny@massey.ac.nz), has made available through his Farside Institute three programs, Hadtree, Prepare, and Trees. These run on DOS systems, and compute bipartition spectra by Hadamard transformations (conjugations and the distance Hadamard), character weighting, distance transformations (including LogDet), base composition tests, resampling schemes, and tree selection. The programs are available from the Farside Institute downloads page at http://imbs.massey.ac.nz/Research/MolEvol/Farside/programs.htm.


David Swofford, of the Laboratory of Molecular Systematics of the Smithsonian Institution, Washington, D.C., has written Freqpars. It implements parsimony analysis based on gene frequencies. The method was described by D. L. Swofford and S. H. Berlocher in a paper in Systematic Zoology 36: 293-325, 1987. The program is available in FORTRAN 77 source code. The search for most parsimonious trees under Swofford and Berlocher's criterion is not very extensive, Swofford notes, because the individual tree evaluations are computationally difficult. The source code in FORTRAN, with documentation, is available by anonymous ftp from onyx.si.edu in directory freqpars.


Jotun Hein, (Institute of Genetics and Ecology, University of Aarhus, 8000 Aarhus C, Denmark) has produced TreeAlign, a multiple sequence alignment program that builds trees as it aligns DNA or protein sequences. It uses a combination of distance matrix and approximate parsimony methods. TreeAlign uses too much memory for it to run on DOS or Macintosh systems but is really designed for a workstation or mainframe. It is available by anonymous ftp at the European Bioinformatics Institute molecular biology software distribution site ftp.ebi.ac.uk in directories pub/software/unix and pub/software/vms.


Another multisequence alignment program that estimates trees as it aligns multiple sequences is ClustalW. Currently it is in version 1.8. It is distributed as C source code and as executables for DOS, Macintosh, and some Unix systems. ClustalW was written by Des Higgins (now at University College, Cork, Ireland) (des@chah.ucc.ie), Julie Thompson (julie@IGBMC.u-strasbg.fr), Toby Gibson, (Gibson@EMBL-Heidelberg.DE), and François Jeanmougin (pingouin@igbmc.u-strasbg.fr). It is a complete rewrite and upgrade of the Clustal and ClustalV packages. which were developed by Des Higgins. New features include the ability to detect read different input formats (NBRF/PIR, Fasta, EMBL/Swissprot); align old alignments; produce phylogenetic trees after alignment (Neighbor Joining trees with a bootstrap option); write different alignment formats (Clustal, NBRF/PIR, GCG, PHYLIP); full command line interface. It is described in the following papers:

ClustalW is available in a number of forms and places:


Ward Wheeler and David Gladstein (wheeler@amnh.org) have written MALIGN, version 2.7, a parsimony-based alignment program for molecular sequences. It implements the original suggestion by Sankoff, Morel, and Cedergren (1973) that alignment and phylogenies could be done at the same time by finding that tree that minizes the total alignment score along the tree. Jotun Hein's program TreeAlign (mentioned above) is another, more approximate but possibly faster, attempt to implement the Sankoff-Morel-Cedergren suggestion. MALIGN is one of the only non-approximate implementations of the original method (Wheeler and Gladstein's other program POY is the other). MALIGN is described in a paper: Wheeler, W. 1996. Optimization alignment: the end of multiple sequence alignment in phylogenetics? Cladistics 12:1-9. MALIGN is available by ftp from the American Museum of Natural History's anonymous ftp site, ftp.amnh.org, in directory pub/molecular. It is available as C source code and as binaries for DOS (with DOS extender), Sun, SGI, HPUX, and Linux. The C source code archive contains as well a special makefile for a version for parallel computation.


[DNASIS icon] Hitachi Software Engineering Co. Ltd. and Molecular Biology Insights, Inc. of Cascade, Colorado, sell DNASIS, a general-purpose DNA and protein sequence analysis system. It has many functions including primer design, plasmid maps, contig assembly, alignment, database searching, and many kinds of protein plots. For our purposes what is relevant is the ability to do multiple sequence alignment by the Higgins-Sharp method of progressive seqeunce alignment (the one used in ClustalV), with one of the results being a UPGMA tree based on pairwise sequence alignment scores. DNASIS is available as Macintosh executables (MacDNASIS version 3.7) and as DNASIS for Windows version 2.6. A description and demo versions are available at its Hitachi web site at http://www.hitachi-soft.com/gs/dnasis/index.htm and at its MBI web site at http://oligo.net/dnasis.htm. Both versions cost $1,895 or $3,000 for a 1-10 user network license.


[GeneDoc icon here] Karl Nicholas (ketchup@cris.com) and Hugh Nicholas (nicholas@psc.edu) of Pittsburgh Supercomputing Center have written GeneDoc, version 2.5, a program for the shading and editing of multiple sequence alignments. Its reads .MSF files and Fasta Files. The alignment can be edited by changing the position of residues in the sequences. GeneDoc includes scoring functions to assist in determining whether your aligment changes are improving the score. Support for obtaining a score via sum-of-pairs or by a phylogenetic tree is included. Phylogenetic trees can be built with either the GUI interface or imported Nexus or Phylip format tree descriptions. The program runs on Windows 3.1, Windows95, and Windows/NT, as both 16-bit and 32-bit executables are distributed. It can be downloaded from its Web site at http://www.cris.com/~ketchup/genedoc.shtml. A Windows NT version for Digital Alpha processors is available from Russell Malmberg at the Botany Department of the University of Georgia (russell@dogwood.botany.uga.edu) by World Wide Web at http://dogwood.botany.uga.edu/malmberg/software.html


Feng Liu and Tao Jiang (jiang@church.dcss.McMaster.CA), of the Department of Computing and Software at McMaster University, Hamilton, Ontario, have written TAAR (Tree Alignment And Reconstruction), version 1.0, which constructs multiple sequence alignment and phylogenies based on the idea of tree alignment. It is a graphical environment capable of "approximately optimal" parsimony-based tree alignment. It can also infer trees by parsimony. It can handle DNA or protein data. It is available as C source code for Unix with X Windows (X11R5) and MOTIF 1.2. It is also available in a version for Linux with Lesstif. It is distributed through its home page at http://www.dcss.mcmaster.ca/~fliu/taar_download.html.


David States (states@ibc.wustl.edu) of the Institute for Biomedical Computing, Washington University, St. Louis, Missouri, has released Ctree version 1.0., a tree alignment program that uses a Hidden Markov Model method of representing the ambiguities in alignments of groups of sequences. The Ctree program is based on a neighbor joining algorithm in which sequences and groups of sequences are represented by Hidden Markov Models. HMMs are aligned using a Smith/Waterman dynamic programming algorithm to find the best local alignment. At each step in building the MSA alignment tree, the highest scoring pair of HMMs are merged into a new HMM. In this sense it is similar to the progressive alignment algorithm used in ClustalW, but the use of an HMM to represent clusters retains more information about the ambiguities than the Clustal algorithm does. It is possible to write the tree to a dendrogram output file.

Ctree is available by anonymous ftp from www.ibc.wustl.edu in directory pub/ctree. It is available as C source code and also as executables for Solaris, SGI, Linux, and Windows95/NT. Ctree is also useable as a server but that version seems not to give trees as output.


Ward Wheeler and David Gladstein (wheeler@amnh.org) of the American Museum of Natural History, New York, have written POY, a program that implements David Sankoff's method of searching for the tree that minimizes a parsimony criterion that includes penalties for gaps, accomplishing both searching for phylogenies and alignments. POY has algorithmic improvements by Wheeler and Gladstein that speed up the algorithm. (Their program MALIGN is the only other program carrying out the full Sankoff proposal). The method is described in a paper: Wheeler, W. 1996. Optimization alignment: the end of multiple sequence alignment in phylogenetics? Cladistics 12:1-9. POY is available in C source or in executables for Linux, HPUX, SGI, Sun, and DOS. It is distributed by ftp from ftp.amnh.org in directory pub/people/wheeler/poy.


Russell Doolittle (rdoolittle@ucsd.edu) and Dafei Feng, of the Department of Chemistry and Biochemistry of the University of California at San Diego, released ALIGN in 1990. A version for Macintoshes was coded by Peter Markeiwicz. ALIGN implements the "progressive alignment" strategy described in their paper: Feng, D.-F. and R. F. Doolittle. 1987. Progressive sequence aligment as a prerequisite to correct phylogenetic trees. Journal of Molecular Evolution 25: 351-360. This is also the basis for the Clustal family of programs as well as the Pileup program in the GCG package. The ALIGN program can align as well as print out a tree (which does not have branch lengths). It uses Doolittle's own formats, and so three other programs are included with ALIGN to convert formats. The programs are distributed by ftp from the EBI ftp software server at ftp.ebi.ac.uk in directory pub/software/mac as file align.hqx.


[PASSML icon here] Pietro Lio, of the Department of Genetics, University of Cambridge (P.Lio@gen.cam.ac.uk), has written PASSML and PASSML_TM, which use likelihood methods with Hidden Markov models to infer phylogeny and also secondary structure from protein data. PASSML is for general proteins and PASSML_TM is for membrane proteins. The methods used are described in the papers: Goldman, N., J. L. Thorne, and D. T. Jones. 1998. Assessing the impact of secondary structure and solvent accessibility on protein evolution. Genetics 149: 445-458, PASSML is described in the paper: Lio, P., N. Goldman, J. L. Thorne and D. T. Jones. 1998. PASSML: combining evolutionary inference and protein secondary structure prediction. Bioinformatics 14: 726-733, and PASSML_TM is described in the paper: Lio, P. and N. Goldman. 1999 Using protein structural information in evolutionary inference: transmembrane proteins. Molecular Biology and Evolution 16: 1696-1710. The programs are available as ANSI C for Unix workstations; PASSML has also been successfully compiled under a port of gcc to a PC. The source code is available via the web page at http://ng-dec1.gen.cam.ac.uk/hmm/Passml.html.


[COMPONENT icon here] Rod Page (dpage@udcf.gla.ac.uk) of the Division of Environmental and Evolutionary Biology of the University of Glasgow has written COMPONENT version 2.0, a program for Windows systems for comparing cladograms for use in phylogeny and biogeography studies. It has many tree comparison and consensus methods, and far more features for biogeographic studies (such as comparing species and area cladograms) than any other package. It also can generate random trees. It runs under Windows 3.0 or higher. Its cost is 40 pounds U.K., and it can be ordered from Anna Hutson at the Department of Botany, Natural History Museum, Cromwell Road, London, SW7 5BD, U.K. (annah@nhm.ac.uk) (fax (0171) 938 9260). Details on how to order may be found at the order form web page at http://taxonomy.zoology.gla.ac.uk/rod/order.html. There is a review of the program in Cladistics 9: 351-353 (1993). COMPONENT has a web site at http://taxonomy.zoology.gla.ac.uk/rod/cpw.html. The documentation is available in Adobe Acrobat at http://taxonomy.zoology.gla.ac.uk/rod/cplite/Manual.html. A very early development Macintosh version ("COMPONENT Lite") is available free from the COMPONENT Lite web site at http://taxonomy.zoology.gla.ac.uk/rod/cplite/guide.html (though Rod says "be prepared for bugs").


[TREEMAP icon here] Rod Page(dpage@udcf.gla.ac.uk), of the Division of Environmental and Evolutionary Biology of the University of Glasgow has written TREEMAP, a free, experimental program for comparing host and parasite phylogenies. It allows you to interactively compare host and parasite trees, construct reconstructions of the history of the association, and perform some simple randomisation tests of hypotheses of cospeciation. The program is available as an executable for Macintoshes or an executable for Windows PCs (the two versions are essentially identical). They can be downloaded from its WWW site: http://taxonomy.zoology.gla.ac.uk/rod/treemap.html. The site also has an online manual, or you can download the documentation as a Postscript file. For a description of the method used by TreeMap, see Page, R.D.M. 1994. Parallel Phylogenies: Reconstructing the history of host-parasite assemblages. Cladistics 10: 155-173.


To next section of software page