| ★ wanayoo — archive 1999 http://evolution.genetics.washington.edu/phylip/software.etc1.html | Nouvelle recherche | Portail wanayoo |
To previous part of Software page
Jun Adachi and Masami Hasegawa have written a package MOLPHY 2.2,
carrying out maximum likelihood inference of phylogenies for either nucleotide
sequences or protein sequences. Their protein sequence maximum likelihood
program, ProtML, is a successor to the one they made available to me for
distribution on a nonsupported basis in PHYLIP, and is much improved over that.
It is one of two protein maximum likelihood programs available. The package is
distributed free in C source code, with documentation,
by ftp from
sunmh.ism.ac.jp.
An executable version for Windows95 or Windows NT on Intel processors, and
also one that works on Windows NT on DEC Alpha processors, is
available from Russell Malmberg at the Botany Department of the
University of Georgia (russell@dogwood.botany.uga.edu)
by World Wide Web
at http://dogwood.botany.uga.edu/malmberg/software.html
Gary Olsen, of the Department of Microbiology, University of Illinois, Urbana, Illinois (gary@phylo.life.uiuc.edu) has developed a speeded-up replacement for my program DNAML coded in C, called fastDNAml. It achieves a number of economies and also is organized so that it can be run on parallel processors -- he and his co-workers have constructed trees of very large size on a high-speed parallel processor. The program can be compiled using the "p4" portable parallel processing toolkit. It can also be run in ordinary serial mode on workstations where it is faster than DNAML.
http://www.cme.msu.edu/RDP/cgis/aftpdir_show.cgi?ftpdir=pub/RDP/programs/fastDNAml&title=Phylogenetic%20tree%20inference%20(fastDNAml;%20C)&showdir=yes
rdp.life.uiuc.edu in directory pub/RDP/programs/fastDNAml.
ftp.bio.indiana.edu
in directory molbio/evolve.
http://www.debian.org/Packages/unstable/misc/fastdnaml.html.
Denis Beaumont
(beaumont@transpac.atlas.fr) has made a parallelized version
of fastDNAml called
VeryfastDNAml. It is parallelized with the TreadMarks
distributed shared memory system, which is a not-quite-free environment for
parallelization that runs on many workstation-class machines. The
C source code of VeryfastDNAml is available by ftp from the
Institut Pasteur server ftp.pasteur.fr in directory
/pub/GenSoft/unix/evolution/FastDNAml as file
fastDNAml-tmk.tar.gz. There is a
web page access to this ftp distribution at
http://bioweb.pasteur.fr/seqanal/soft-pasteur.html#veryfastdnaml,
which includes a link to the TreadMarks project.
Ziheng Yang of the Department of
Genetics and Biometry, University College London,
(z.yang@ucl.ac.uk) has released
PAML, version 2.0g, a package of programs for the maximum
likelihood analysis of nucleotide or protein sequences, including codon-based
methods that take into account both amino acids and nucleotides.
The programs can estimate branch lengths in a phylogenetic tree and parameters
in the evolutionary model such as the transition/transversion rate ratio, the
gamma parameter for variable substitution rates among sites, rate
parameters for different genes, and synonymous and nonsynonymous substitution
rates. They can also test evolutionary models, calculate
substitution rates at particular sites, reconstruct ancestral nucleotide or
amino acid sequences, simulate DNA and protein sequence evolution,
compute distances based on the synonymous and nonsynonymous changes,
and of course do phylogenetic tree reconstruction by
maximum likelihood and Bayesian Markov Chain Monte Carlo methods. The
strength of the package lies in its rich implementation of
evolutionary models, though Yang coments that
tree-making is not a strong point of the current version.
The autocorrelated gamma distribution implementation is analogous to
the Hidden Markov Model scheme available in PHYLIP. The package is
available as
ANSI C source code for Unix systems, as PowerMac executables and
as executables that run on Windows95, Windows98, and WindowsNT.
See the PAML
web page at http://abacus.gene.ucl.ac.uk/ziheng/paml.html.
It can be downloaded by ftp from abacus.gene.ucl.ac.uk
in directory pub/paml.
Bret Larget and Donald Simon of the Department of
Mathematics and Computer Science, Duquesne University, Pittsburgh,
Pennsylvania (bambe@mathcs.duq.edu) have written
BAMBE (Bayesian Analysis in Molecular Biology and
Evolution) version 2.01 beta, a program for Bayesian analysis of phylogenies
with DNA sequence data. It uses a prior distribution of trees
and arearrangement mechanism introduced in the paper:
Mau, B., M. A. Newton, and B. Larget. 1997. Bayesian phylogenetic inference
via Markov chain Monte Carlo methods. Molecular Biology and
Evolution 14: 717-724.
The trees and parameter values are sampled by a Metropolis
algorithm Markov Chain Monte Carlo sampling.
The resulting posterior
distribution can be used to characterize the uncertainty about not
only the tree, but the parameters of the substitution model as
well.
The program is in C++ source code for Unix, and is distributed from
its web page at
http://www.mathcs.duq.edu/larget/bambe.html.
David Posada and Keith Crandall
at the Department of Biology, Brigham Young University
(dp47@email.byu.edu) has released
Modeltest version 3.0, a program to test a hierarchy
of statistical models of DNA evolution using the Likelihood Ratio Test
criterion and the AIC (Akaike Information Criterion). The likelihood
values are obtained by running PAUP*.
MODELTEST accepts likelihood scores corresponding to 56 models of DNA
substitution including whether transition and transversion rates are
equal, whether rates at different sites are equal, and whether there are
invariant sites. Modeltest is described in the paper:
Posada, D. and K. A. Crandall. 1998. MODELTEST: testing the model of DNA
substitution. Bioinformatics 14: 817-818.
It is available as executables for Macintosh and PowerMac, for Windows95/98/NT,
for Linux and for Suns, and in source code as C for the Metrowerks C compiler.
It is distributed from
its web
site at
http://bioag.byu.edu/zoology/crandall_lab/modeltest.htm.
Nick Grassly,
currently of the Zoologisches Institut, Universität München
(grassly@zi.biologie.uni-muenchen.de), has written PLATO,
version 2.11, a program that takes sequential PHYLIP-style DNA sequences
followed by their maximum likelihood phylogeny, and using a likelihood approach
with sliding window analysis and Monte Carlo simulation
of the null distribution detects anomalously evolving regions in the DNA
sequences and assesses their significance.
This may lead to the detection of, for example, recombination, gene conversion
or convergence, or reveal variable selective pressures along the gene sequence.
A general substitution model is used that can allow the test to reveal
differences due to recombination while ignoring those due to varying rate
of evolution.
The method is described in the paper: Grassly, N. C., and E. C. Holmes. 1997.
A likelihood method for the detection of selection and recombination using
sequence data. Molecular Biology and Evolution 14: 239-247.
It is available for Macintoshes (including PowerMacs) or in source code for Unix systems.
It requires substantial amounts off memory, especially when sequences analysed
are long. Use of Power Macintoshes or UNIX systems is recommended.
It is distributed free
from the University of Oxford Zoology
Web server at
http://evolve.zoo.ox.ac.uk/Plato/Plato2.html.
Mika Salminen and Wayne Cobb (msalminen@hiv.hjf.org and wcobb@reed.hjf.org), of the Henry M. Jackson
Foundations for the Advancement of Military Medicine, Walter Reed Army
Institite of Research, Bethesda, Maryland, have released the
Bootscanning Package, version 1.0beta. This is a series
of shell scripts and programs that analyze DNA sequences for evidence of
recombination. It breaks the sequence into separate pieces that are
analyzed for the bootstrap support of various groups, and it looks for
evidence of significant conflict among trees for different parts of the
sequence. The programs are currently available only as Sun executables.
They require GDE 2.2a and
PHYLIP version 3.4 to work.
They are available by anonymous ftp from from
http://www.ktl.fi in directory /hiv/mirrors/pub/programs.
Gráinne McGuire and Frank Wright (grainne@bioss.sari.ac.uk and frank@bioss.sari.ac.uk)
of Biomathematics and Statistics Scotland, in Dundee, have released
TOPAL, which checks for evidence of
past recombination events, by looking for changes in the inferred phylogenetic tree TOPology
between adjacent regions of a multiple sequence ALignment. Their method
detects recombinations by sliding a window along a sequence alignment, and
measuring the discrepancy between the trees suggested by the first and
second halves of the window, using distance matrix methods. This is described
in the paper:
McGuire, G., F. Wright, and M. J. Prentice. 1997. A graphical method for
detecting recombination in phylogenetic data sets. Molecular Biology
and Evolution 14: 1125-1131.
The TOPAL program is also described in a paper:
McGuire, G. and F. Wright. 1997. TOPAL: recombination detection in DNA and
protein sequences. Bioinformatics 14: 219-220.
TOPAL is a set of Unix Bourne shell scripts and C code, plus four programs
in C from my PHYLIP package.
These are available from the TOPAL
web site
at http://www.bioss.sari.ac.uk/~frank/Genetics/topal.html.
Ingrid Jakobsen and Simon Easteal
of Australian National University,
Canberra, have released reticulate.
(Ingrid Jakobsen is now
at the Institute of Molecular Evolutionary Genetics, Pennsylvania State
University, and her e-mail address is ibj1@psu.edu).
It is a compatibility matrix program for DNA sequences that has features
designed to test for evidence of reticulate evolution (such as recombination).
The program computes and displays a pairwise compatibility matrix
for all pairs of sites. It can randomize the order of sites and compute
the fraction of compatible sites in a region for the randomizations, to
test whether there is a pattern suggesting reticulation.
The program is distributed as C source code for Unix and X Windows, though
there are some limited ways of running it without X Windows. It is
described in the paper:
Jakobsen, I. B. and S. Easteal. 1996.
A program for calculating and displaying compatibility matrices as an aid in determining reticulate evolution
in molecular sequences. CABIOS 12: 291-295. It is
available from its
web site at
http://jcsmr.anu.edu.au/dmm/humgen/ingrid/reticulate.htm.
Kim Fisker (kfisker@daimi.aau.dk)
of the Computer Science Department at Aarhus University, Denmark has
released RecPars, which does
a parsimony analysis of DNA sequences. It tries to find
the best phylogenies for different regions of the sequences and
thereby postulating a recombination event between these segments.
The method is described in a paper: Hein, J. 1993. A heuristic method to
reconstruct the history of sequences subject to recombination.
Journal of Molecular Evolution 36: 396-406.
RecPars is available as C source code for Unix. It is distributed
by ftp
from ftp.daimi.aau.dk in directory
pub/empl/kfisker/programs/RecPars.
John Maynard Smith and Noel Smith
of the School of Biological Sciences of the University of Sussex
(noelsmith@yahoo.com) have released programs to
carry out their homoplasy test for recombination in
sequences. The test is described in a paper:
Maynard Smith, J. and N. H. Smith. 1998. Detecting recombination from gene
trees. Molecular Biology and Evolution 15: 590-599.
The programs are distributed in QBASIC for DOS and must be run using
QBASIC. They are available from Maynard Smith's
web site
at http://www.biols.susx.ac.uk/Home/John_Maynard_Smith/.
Andrew Rambaut of the Department of Zoology, University
of Oxford, England (andrew.rambaut@zoo.ox.ac.uk) has
produced LARD (Likelihood Analysis of Recombination in DNA)
version 2.2, a program to detect the presence of recombination in a set
of sequences. LARD looks at the set of sequences to discover which are the
most plausible parents of a potentially recombinant sequence, and performs
a likelihood ratio test for each possible breakpoint position of whether the
three-species tree differs on the two sides of the breakpoint. LARD is
described as an extension of a method suggested by John Maynard Smith:
Maynard Smith, J. 1992. Analysing the mosaic structure of genes.
Journal of Molecular Evolution 34: 126-129. It is
described in a paper: Holmes, E. C., M. Worobey, and A. Rambaut. 1999.
Phylogenetic evidence for recombination in dengue virus. Molecular
Biology and Evolution 16: 405-409. LARD is available
as C source code and as a Macintosh executable from
its web site
at http://evolve.zoo.ox.ac.uk/Lard/Lard.html.
Andrew Rambaut of the Department of
Zoology, University of Oxford, (andrew.rambaut@zoo.ox.ac.uk) and
Nick Grassly, currently of the Zoologisches Institut, Universität
München
(grassly@zi.biologie.uni-muenchen.de),
have written SPOT (Sequence Parameters Of Trees).
SPOT is a program that will calculate the likelihood of a given tree topology
for a set of aligned nucleotide sequences. For each topology, SPOT will
estimate the maximum likelihood values of branch lengths and other parameters
of the model of nucleotide evolution that has been chosen. Such parameters
include the ratio of transitions to transversions (TS/TV ratio) and relative
rates of substitution at different codon positions. Branch lengths can also
be constrained to assume a molecular clock hypothesis. Multiple datasets
and multiple trees can be analysed which is useful for performing Monte
Carlo simulations of hypothesis (parametric bootstraps). Although SPOT does
not estimate tree topology, an accompanying program, SPOTSHELL, will iterate
between fastDNAml and SPOT until the maximum likelihood parameters and topology
has been found (or at least something close to it).
SPOT is available as C source code for Unix workstations, or as Macintosh sources and
executables.
It can be obtained from the
SPOT Web page
at http://evolve.zoo.ox.ac.uk/Spot/Spot.html.
Gary Olsen of the Department of Microbiology, University
of Illinois, Urbana, Illinois (gary@phylo.life.uiuc.edu) has written
dnarates version 1.0. It reads a set of DNA sequences
and a tree, and for that tree makes a maximum likelihood estimate of the
rate of evolution at each site. This is done by taking the rate at each
site as a separate parameter and maximizing the likelihood with respect to
all those parameters. The program is available as generic C source code.
It is based in part (with my permission) on code from my PHYLIP program DNAML. dnarates is available
by ftp
from the IUBIO ftp server at
ftp://rdp.life.uiuc.edu/pub/RDP/programs/DNArates/.
Mike Charleston (mcharles@udcf.gla.ac.uk)
of the Division of Environmental and Evolutionary Biology of the University
of Glasgow
has developed Spectrum, a program for finding bipartition spectra
from phylogenetic
molecular and distance data, according to the method of Hendy et al.
(1994) (Hadamard transforms)
for moderately sized data sets (up to 18 taxa). The program also
implements a
branch-and-bound search for the "closest tree" - that is, the tree whose
expected spectrum is closest to the spectrum derived from the observed
PowerMac, 68k Macintosh, and Windows95 or Windows NT executables are
available from its
Web site in the Glasgow Taxonomy web
pages:
http://taxonomy.zoology.gla.ac.uk/~mac/spectrum/spectrum.html.
Ingrid Jakobsen, Susan Wilson, and Simon Easteal,
of Australian National University,
Canberra, have released partimatrix. (Ingrid Jakobsen is now
at the Institute of Molecular Evolutionary Genetics, Pennsylvania State
University, and her e-mail address is ibj1@psu.edu). This program
computes a "partition matrix" from aligned DNA sequence data. The method
finds partitions of the sequences into two groups and presents a matrix
which describes the conflict and agreement among these partitions. The
objective is to discover parts of the DNA sequence which imply different
trees. It is described in the paper
by I. B. Jakobsen, S. R. Wilson and S. Easteal. 1997.
The Partition Matrix: Exploring variable phylogenetic signals along nucleotide sequence alignments.
Molecular Biology and Evolution 14: 474-484.
The program is distributed as C source code for Unix systems with X Windows.
It is available from
its web site at
http://jcsmr.anu.edu.au/dmm/humgen/ingrid/partimatrix.htm
.
Pablo Goloboff, of INSUE - Fundación e Instituto
Miguel Lillo 205, 4000 S. M. de
Tucumán, Argentina, has written
Nona (Noname), PiWe
(Parsimony with Implied WEights),
and SPA to carry out parsimony including weighted
parsimony analyses. Nona searches for most parsimonious trees according to
character weights defined by the user a priori. Pee-Wee calculates weights of
the characters by a method introduced by Goloboff, a
noniterative version of J. S. Farris's "successive weighting". It was described
in Goloboff's paper in Cladistics 9: 83-91, 1993.
SPA is a generalized parsimony program that allows differential weighting of
changes between different states.
Nona is said to be faster than other parsimony programs.
A Windows 95/98/NT version of Nona which includes the functionality of Piwe
and SPA is available as shareware (with a free 30-day
trial period) from
its web page at
http://www.cladistics.com/about_nona.htm. The shareware fee of
$40 should be paid to the author at the above
address or to James M. Carpenter,
Department of Entomology, American Museum of Natural History, Central Park West at
79th Street, New York, NY 10024. Send the money and the name in which the
copies are to be registered. An earlier demo version of these programs
which runs on DOS is also
available. It requires the user to
hit an extra key each time they execute a command.
It is can be fetched by anonymous from ftp.vims.edu
in directory other/hennig as file pars-pag.exe, a self-extracting
DOS Zip file.
Yasuo Ina of the National Institute of Agrobiological
Resources, Tsukuba, Japan
ha developed ODEN, a package of programs for doing
distance matrix analyses on nucleotide or protein sequences. It is described
in CABIOS 10: 11-12 (1994). It is available free
by anonymous ftp from
directory pub/unix/oden on ftp.dna.affrc.go.jp as C source code for Unix systems.
A. Luettke and R. Fuchs have written MacT, a package of programs for
Macintoshes that compute distances and compute Neighbor-Joining phylogenies for
them. The programs work on 4 through 26 sequences, and source code in
Microsoft QuickBasic is provided as well as compiled executables. The package
is free and is available on the molecular biology software servers.
For example, it is available on by anonymous ftp on the Indiana University IUBIO server
ftp.bio.indiana.edu it will be found in directory molbio/mac. The programs are
described in CABIOS 8: 591-594, 1992.
Andrey A. Zharkikh, Andrey Rzhetsky, and co-workers in the
Institute of Cytology and Genetics, Siberian Branch of the Russian Academy of
Sciences, Novosibirsk, Russia, have produced VOSTORG, a package of
programs for alignment (both manual and automatic) and inferring phylogenies by
distance methods and parsimony for molecular sequences.
(Zharkikh and Rzhetsky are currently in the US; their e-mail addresses are
zharkikh@myriad.com and andrey@genome2.cpmc.columbia.edu). VOSTORG runs on under DOS on
PC-compatibles and includes some rather fancy graphics (for DOS).
It is available from
its Web page in Russia
from http://molevol.bionet.nsc.ru/vs.htm.
The programs are described in a paper: Zharkikh, A. A., A-Yu. Rzhetsky,
P. S. Morosov, T. L. Sitnikova, and J. S. Krushkal. 1991. VOSTORG: a package of
microcomputer programs for sequence analysis and construction of phylogenetic
trees. Gene 101: 251-254.
Walter Fitch (wfitch@uci.edu), of the Department
of Ecology and Evolutionary Biology, of the University of California at
Irvine, has available by anonymous ftp at daedalus.bio.uci.edu in
directory pub/outgoing/evoprog about 20 programs
which carry out various kinds of phylogeny estimation and related tasks.
They are available in source code in FORTRAN 77,
(except for a few which are in C) and also as Sun SPARC executables and as
DOS executables. They include:
daedalus.bio.uci.edu in directory
pub/outgoing/evoprog.
There is also TDRAW which draws a tree in Postscript. This
program is in C, and is not available as a DOS executable.
It is available
in directory pub/outgoing/tdraw.
Nicholas Galtier of the University of Lyon (galtier@biomserv.univ-lyon1.fr)
has written Phylo_win, a "graphic interface" for molecular
phylogenetic inference. It performs neighbor-joining, parsimony and
maximum likelihood methods and can bootstrap with any of them. Many distances
can be used including Jukes & Cantor, Kimura, Tajima & Nei, Galtier & Gouy
(1995), LogDet for nucleotidic sequences, Poisson correction for protein
sequences, Ka and Ks for codon sequences. Species and sites to include in the
analysis are selected by mouse. Reconstructed trees can be drawn, edited,
printed, stored, evaluated according to numerous criteria.
Taxonomic species groups and sets of conserved regions can be defined by
mouse in both tools and stored into sequence files, thus avoiding multiple
data files. It is entirely mouse-driven. Most usual sequence file formats are
read: CLUSTAL, FASTA, PHYLIP, MASE. It runs under X windows on many Unix
workstations.
It is described in the paper:
Galtier, N., M. Gouy, and C. Gautier. 1996. SeaView and Phylo_win, two graphic
tools for sequence alignment and molecular phylogeny. Computer Applications
in the Biosciences 12: 543-548.
It is distributed as C source code (to compile it one needs the NCBI Vibrant
tool kit). It is also available
as executables for SunOS, Solaris, SGI Unix,
IBM RISC Unix, Linux, HP/UX, and DEC Alpha (Digital Unix). It can be
fetched from
its web page at http://pbil.univ-lyon1.fr/software/phylowin.html.
It can also be obtained by anonymous ftp from
biom3.univ-lyon1.fr in directory pub/mol_phylogeny.
A PC Linux executable
is available at http://evolution.bmc.uu.se/~thomas/mol_linux.
A Digital OpenVMS executable is
also available
as http://seqaxp.bio.caltech.edu:8000/pub/SOFTWARE/phylo_win_vms.zip.
sales@exetersoftware.com. Their
toll-free telephone number is 800-842-5892, their not-so-free
phone number is +1-631-689-7838, and their fax number is +1-631-689-0103.
Their mailing address is 100 North Country Road, Setauket, NY 11733-1345 USA.
FAX, or phone (toll-free telephone within the USA).
Further information is available on their
Web page
at http://www.exetersoftware.com/cat/ntsyspc.html.
Rino Zandee (zandee@rulsfb.leidenuniv.nl),
of the Institute of Evolutionary and Ecological Science, Van der Klaauw
Laboratory, Leiden University, has written CAFCA version 1.5j,
the Collection of APL Functions for Comparative Analysis. It carries out a
search for the most parsimonious tree with discrete-character data (either
two-state or multistate), using a search for cliques of component
compatibility (monothetic subsets) to propose the candidates for most
parsimonious trees. The program is written as functions in the APL language,
but Macintosh and PowerMac executables are distributed. The program is free
and is available from the
CAFCA Web Site
http://wwwbio.leidenuniv.nl/~zandee/cafca.html.
Korbinian Strimmer
http://www.debian.org/Packages/unstable/misc/puzzle.html.
Mike Holder (holder@mbl.edu) and Andrew Roger(roger@mbl.edu) of the Marine Biological Laboratory in
Woods Hole, Massachusetts are distributing a shell script program for
Unix systems, puzzleboot that allows the analysis of
multiple data sets with TREE-PUZZLE. It is
designed for use with the distance matrix option of TREE-PUZZLE, to make use of
the distance calculation methods.
It is available from the TREE-PUZZLE
web page at http://www.tree-puzzle.de.
kns@phy.auckland.ac.nz) of the
Department of Physics of the University of Auckland, New Zealand has released
version 1.0 of STATGEOM.
It carries out computation of the statistical geometry in distance and in
sequence space of a set of aligned DNA/RNA, amino acid or binary sequences.
The user can decide to
either compute the overall tree-likeness of the whole set, or a
certain subset, or given a tree of the sequences compute the
reliability of certain edges in the tree.
Postscript files of the graphs of the
statistical geometry are automatically generated.
A sequence reformatting utility allow various sequence formats to be
read in.
STATGEOM is written in ANSI C; source code with documentation and a Sun SPARC executable are
available by anonymous ftp at cage.mpibpc.gwdg.de (or 134.76.209.64) in directory
pub/kniesel.
The method of statistical geometry was originally published in:
Eigen, M., Winkler-Oswatitsch, R. and Dress, A. 1988.
Statistical geometry in sequence space: a method of comparative
sequence analysis. Proc. Natl. Acad. Sci. USA 85: 5913-5917.
Rainer Wetzel and Daniel Huson of the University
of Bielefeld (huson@mathematik.uni-bielefeld.de)
have developed a program SplitsTree2 for
SplitsTree2 carries out the split decomposition method of A. Bandelt and A. Dress,
Bandelt and Dress's p-splits, and spectral analysis (Hendy, Penny, Szekely,
Steel & Erdos). It can process sequence or restriction site data, and
can do does bootstrapping. It also contains an implementation of
Cooper, Penny and Steel's method for dating divergences and also their
molecular clock test.
(Molecular Phylogenetics 1: 242-252 (1992)).
It is also discussed in the paper: Huson, D. H. 1998. SplitsTree: analyzing
and visualizing evolutionary data. Bioinformatics 14: 68-73.
SplitsTree2 is currently available as a Mac program or in a Unix version for a number of different
machines (Sun, SGI, DEC and HP). These use the program ghostview to draw the computed
graph. The Mac version draws the graph in its own window, and the picture can
be copied and pasted or printed in the usual way. There is no Windows
version at present. It is available by ftp from
ftp.uni-bielefeld.de in directory pub/math/splits.
SplitsTree2 is software under development.
A server is also maintained
which uses SplitsTree 2 to analyze data submitted via its web page.
Igor Kuznetov and Pavel Morozov
(kuznets@mailhost.bionet.nsk.su) of the Institute of Cytology
and Genetics, Novosibirsk, Russia, have produced GEOMETRY,
a package for nucleotide sequence analysis using the method of
statistical geometry in sequence space (M. Eigen, R. Winkler-Oswatitsch, and
A. Dress. 1988. Statistical geometry in sequence space: A method of quantitative
comparative sequence analysis, Proc. Natl. Acad. Sci. USA 85: 5913-5917). The program is described in an article by Kuznetsov and
Morozov in 1996 in CABIOS 12: 297-301.
The package uses the same data formats for sequence and tree input as
the ones used in VOSTORG package.
GEOMETRY is available as a DOS executable.
It is available for downloading
by Web
from http://molevol.bionet.nsc.ru/soft.htm
or by ftp from ftp.bionet.nsk.su
in directory incoming/molevol and also from the
EMBL file server ftp.ebi.ac.uk in directory
pub/software/dos.
Vincent Berry of the Université Jean Monnet in
St.-Etienne, France (vberry@univ-st-etienne.fr) has released
PhyloQuart version 1.3, a package of programs inferring
phylogenies from quartets. It is able to use either nucleotide sequences or
distances. It implements the Q* method of tree reconstruction, which is
inspired by the work of Bandelt and Dress, and is described in the
forthcoming paper: Berry, V. and O. Gascuel. Inferring Evolutionary Trees with
Strong Combinatorial Evidence. Theoretical Computer Science, to
appear. PhyloQuart is available as C source code which can be compiled on
Unix systems, from
its web site at
http://www.univ-st-etienne.fr/eurise/LOGICIELS/PHYLOQUART/main-pq.html.
PhyloQuart is also available as a Web server from the server of the Institut Pasteur.
Stephen J. Willson (swillson@iastate.edu) of the Department of Mathematics, Iowa State University, has produced a package of programs to infer phylogenies from quartets of species. They infer phylogenies of individual quartets by parsimony, and in combining them use information on how strongly the phylogeny for that quartet is preferred over its alternatives, or by measures of how well the group fits into a given placement on a tree, as judged by quartets. The methods are described in two papers: Willson, S. J. 1998. Measuring inconsistency in phylogenetic trees, Journal of Theoretical Biology 190: 15-36, and Willson, S. J. 1998. Building phylogenetic trees from quartets by using local inconsistency measures . Molecular Biology and Evolution 16: 685-693. The programs are in C and are described as having successfully been compiled on PowerMac systems using the Codewarrior C compiler. PowerMac executables are also provided. The programs are available at Willson's software web site at http://www.public.iastate.edu/~swillson/software.html.
James Lake of the Department of Molecular, Cell and
Developmental Biology of the University of California, Los Angeles
(lake@mbi.ucla.edu) has released Gambit, which
implements a method called Boostrapper's Gambit. The method involves
bootstrap sampling sequences, computing trees for quartets of species, and
assembling larger trees out of quartets that have significant boostrap
support. One of the methods available to estimate trees from
the quartets is paralinear (LogDet) distances. Other distance methods and
parsimony are also available. The program is available as a DOS executable,
free to noncommercial users on a trial basis until January 15, 2001.
Commercial users are asked to pay $50 on a shareware basis.
The program is available at
its web site at
http://www.lifesci.ucla.edu/mcdbio/Faculty/Lake/Research/Programs/.
Arne Röhl, Peter Forster, and Hans-Jürgen Bandelt (Forster is at pf223@cus.cam.ac.uk) have written Network 2.0b, a program to infer networks (which have more connections than trees).
The networks are median-joining networks, a method which is described
in a paper: Bandelt, H-J., P. Forster, and A. Röhl. 1999. Median-joining
networks for inferring intraspecific phylogenies. Molecular Biology
and Evolution 15: 1108-1114. The program is available as shareware
(free until 1 June 2000) as a DOS executable from Fluxus Engineering at
its web site at
http://www.fluxus-engineering.com/sharenet.htm.
Pierre Rioux and Tim Littlejohn of the Informatics Division of the
Organelle Genome Megasequencing Program at the Universite de Montreal has made
Available PARBOOT, a program that takes bootstrap sampled data sets and splits
them up, submitting each to a different computer, so as to run bootstrapping
quickly on networks of computers. It is available free as C source code
by ftp
from megasun.bch.umontreal.ca in directory pub/parboot. It requires a
networked system of computers with PHYLIP, a Perl interpreter, and
appropriate accounts and permissions.
Naoko Takezaki (ntakezak@lab.nig.ac.jp)
of the Center for Information Biology of the National Institute of Genetics,
Mishima, Japan, has written Lintre (Phylogenetic tests of
the molecular clock and linearized tree), a package of programs for
Sun workstations. The programs include:
http://cib.nig.ac.jp/dda/ntakezak.html and
by anonymous
ftp from the IUBio
server ftp.bio.indiana.edu in directory molbio/evolve.
Naoko Takezaki (ntakezak@lab.nig.ac.jp)
of the Center for Information Biology, National Institute of Genetics, Mishima, Japan, has written
njbafd which constructs a neighbor-joining tree or a UPGMA
tree from microsatellite data and other allele frequency data. Bootstrapping
can be carried out. The program includes Goldstein et al.'s distance for
microsatellite loci. There are is a source code Unix version and an executable
DOS version. They are available
from her web site at
http://cib.nig.ac.jp/dda/ntakezak.html.
They are
also available by ftp in directory molbio/evolve of the
IUBIO archive.
Joaquin Dopazo of
the Bioinformatics department of GlaxoWellcome SA, Spain
(jd19662@glaxowellcome.co.uk) has written
WET (Windows Easy Tree), version 1.3, which is an easy-to-use program
for inferring phylogenies from sequence data by distance matrix methods.
The main goal in the development of WET was to make a really user friendly
program able to interact with other phylogenetic packages.
WET can import files of a number of different formats. It
calculates distances by a number of different methods and constructs
phylogenetic trees using neighbor-joining, UPGMA and WPGMA procedures. It is a
Windows95/98/NT executable. It is
available from
its web site at http://www.cnb.uam.es/~dopazo/software/wet.html.
David Penny (Institute of Molecular Biosciences, Massey University,
Palmerston North, New Zealand) has been offering for free distribution two
DOS programs, one a fast parsimony program, TurboTree. There is also another,
Great Deluge, an approximate
search for the most parsimonious tree by a quasi-random method. He tells me
that funding exigiencies are such that he may soon have to start charging for
these. His electronic mail address is dpenny@massey.ac.nz.
David Penny of the Institute of Molecular Biosciences,
Massey University, Palmerston North, New Zealand (dpenny@massey.ac.nz),
has made available through his Farside Institute three programs, Hadtree,
Prepare, and Trees. These run on DOS
systems, and compute bipartition spectra by Hadamard transformations (conjugations and the
distance Hadamard), character weighting, distance transformations (including
LogDet), base composition tests, resampling schemes, and tree selection.
The programs are available from the
Farside Institute downloads page at
http://imbs.massey.ac.nz/Research/MolEvol/Farside/programs.htm.
David Swofford, of the Laboratory of Molecular
Systematics of the Smithsonian Institution, Washington, D.C., has written
Freqpars. It implements parsimony analysis based on gene
frequencies. The method was described by D. L. Swofford and S. H. Berlocher
in a paper in Systematic Zoology 36: 293-325, 1987. The program
is available in FORTRAN 77 source code. The search for most parsimonious
trees under Swofford and Berlocher's criterion is not very extensive,
Swofford notes,
because the individual tree evaluations are computationally difficult.
The source code in FORTRAN, with documentation, is available by
anonymous ftp from onyx.si.edu in directory freqpars.
Jotun Hein, (Institute of Genetics and Ecology, University of Aarhus,
8000 Aarhus C, Denmark) has produced TreeAlign, a multiple sequence alignment
program that builds trees as it aligns DNA or protein sequences. It uses a
combination of distance matrix and approximate parsimony methods. TreeAlign
uses too much memory for it to run on DOS or Macintosh systems but is really
designed for a workstation or mainframe. It is available
by anonymous ftp at
the European Bioinformatics Institute molecular biology software
distribution site ftp.ebi.ac.uk
in directories pub/software/unix and pub/software/vms.
Another multisequence alignment program that estimates trees as it
aligns multiple sequences is ClustalW. Currently it is
in version 1.8. It is distributed as C source
code and as
executables for DOS, Macintosh, and some Unix systems.
ClustalW was written by Des Higgins (now at University College, Cork, Ireland)
(des@chah.ucc.ie),
Julie Thompson (julie@IGBMC.u-strasbg.fr), Toby Gibson,
(Gibson@EMBL-Heidelberg.DE), and
François Jeanmougin (pingouin@igbmc.u-strasbg.fr).
It is a complete rewrite and upgrade of the Clustal and ClustalV packages.
which were developed by Des Higgins. New features include the
ability to detect read different input formats (NBRF/PIR, Fasta,
EMBL/Swissprot); align old alignments; produce phylogenetic trees after
alignment (Neighbor Joining trees with a bootstrap option); write different
alignment formats (Clustal, NBRF/PIR, GCG, PHYLIP); full command line
interface. It is described in the following papers:
ftp.bio.indiana.edu
and ftp.ebi.ac.uk. In
the Indiana archive one must enter directory molbio/align,
and in the EBI archive it is in directory pub/software
in four directories unix/clustalw, vms/clustalw, mac/clustalw, and DOS/clustalw. These also contain the older ClustalV executables, as well
as a version, ClustalX, that has a windowing interface.
Clustal X is made available as executables for PowerMac, PC (32 bit), and UNIX
(Linux, Alpha, SGI, Sun).
ftp-igbmc.u-strasbg.fr
in directory pub and there is a
description of ClustalX on
its web
page at http://www-igbmc.u-strasbg.fr/BioInfo/ClustalX/Top.html
and in a paper: Thompson, J. D., T. J. Gibson, F. Plewniak, F. Jeanmougin,
and D. G. Higgins. 1997. The ClustalX windows interface:
flexible strategies for multiple sequence alignment aided by quality analysis
tools. Nucleic Acids Research 24: 4876-4882.
http://www.sgi.com/chembio/resources/clustalw/parallel_clustalw.html.
E-mail contact at SGI for this version is Dmitri Mikhailov
(dmitri@sgi.com).
http://yeamob.pci.chemie.uni-tuebingen.de/Archiv/ClustToTree.html.
Ward Wheeler and David Gladstein (wheeler@amnh.org)
have written MALIGN, version 2.7, a parsimony-based alignment program for molecular sequences. It implements the original
suggestion by Sankoff, Morel, and Cedergren (1973) that alignment and
phylogenies could be done at the same time by finding that tree that minizes
the total alignment score along the tree. Jotun Hein's program TreeAlign
(mentioned above) is another, more approximate but possibly faster, attempt to
implement the Sankoff-Morel-Cedergren suggestion. MALIGN is one of the only
non-approximate implementations of the original method (Wheeler
and Gladstein's other program POY is the other).
MALIGN is described in a paper: Wheeler, W. 1996.
Optimization alignment: the end of multiple sequence alignment in
phylogenetics? Cladistics 12:1-9.
MALIGN is available by ftp from the American Museum of Natural
History's anonymous ftp site, ftp.amnh.org, in directory pub/molecular. It is available as C source code and as binaries for
DOS (with DOS extender), Sun, SGI, HPUX, and Linux. The
C source code archive contains as well a special makefile for a version
for parallel computation.
Hitachi Software Engineering Co. Ltd. and
Molecular Biology Insights, Inc. of Cascade, Colorado, sell
DNASIS, a general-purpose DNA and protein sequence analysis system.
It has many functions including primer design, plasmid maps, contig assembly,
alignment, database searching, and many kinds of protein plots. For our
purposes what is relevant is the ability to do multiple sequence alignment
by the Higgins-Sharp method of progressive seqeunce alignment (the one used
in ClustalV), with one of the results being a UPGMA tree based on pairwise
sequence alignment scores. DNASIS is available as Macintosh executables
(MacDNASIS version 3.7) and as DNASIS for Windows version 2.6. A
description and demo versions are available at
its Hitachi web site at http://www.hitachi-soft.com/gs/dnasis/index.htm and at
its MBI web site
at http://oligo.net/dnasis.htm. Both versions
cost $1,895 or $3,000 for a 1-10 user network license.
Karl Nicholas (ketchup@cris.com) and Hugh Nicholas (nicholas@psc.edu)
of Pittsburgh Supercomputing Center have written GeneDoc,
version 2.5, a program for the shading and
editing of multiple sequence alignments. Its reads .MSF files and Fasta Files.
The alignment can be edited by changing the position of residues in the
sequences. GeneDoc includes scoring functions to assist in determining
whether your aligment changes are improving the score. Support for obtaining
a score via sum-of-pairs or by a phylogenetic tree is included. Phylogenetic
trees can be built with either the GUI interface or imported Nexus or Phylip
format tree descriptions. The program runs on Windows 3.1, Windows95,
and Windows/NT, as both 16-bit and 32-bit executables are distributed.
It can be downloaded from
its Web site at http://www.cris.com/~ketchup/genedoc.shtml.
A Windows NT version for Digital Alpha processors is available from
Russell Malmberg at the Botany Department of the
University of Georgia (russell@dogwood.botany.uga.edu)
by World Wide Web
at http://dogwood.botany.uga.edu/malmberg/software.html
Feng Liu and Tao Jiang (jiang@church.dcss.McMaster.CA), of
the Department of Computing and Software at McMaster University,
Hamilton, Ontario,
have written TAAR (Tree Alignment And Reconstruction),
version 1.0, which constructs multiple
sequence alignment and phylogenies based on the idea of tree alignment.
It is a graphical environment capable of "approximately optimal"
parsimony-based tree alignment. It can also infer trees by parsimony.
It can handle DNA or protein data.
It is available as C source code for Unix with X Windows (X11R5) and MOTIF 1.2.
It is also available in a version for Linux with Lesstif.
It is distributed through
its home page
at http://www.dcss.mcmaster.ca/~fliu/taar_download.html.
David States (states@ibc.wustl.edu)
of the Institute for Biomedical Computing, Washington University, St.
Louis, Missouri, has released
Ctree version 1.0., a tree alignment program that
uses a Hidden Markov Model method of representing the ambiguities in
alignments of groups of sequences.
The Ctree program is based on a neighbor joining algorithm in which
sequences and groups of sequences are represented by Hidden Markov
Models. HMMs are aligned using a Smith/Waterman dynamic programming
algorithm to find the best local alignment. At
each step in building the MSA alignment tree, the highest scoring pair
of HMMs are merged into a new HMM. In this sense it is similar to the
progressive alignment algorithm used in ClustalW, but
the use of an HMM to represent clusters retains more information about the
ambiguities than the Clustal algorithm does. It is possible to write the
tree to a dendrogram output file.
Ctree is available by anonymous ftp
from www.ibc.wustl.edu in
directory pub/ctree. It is available as C source code and also as
executables for Solaris, SGI, Linux, and Windows95/NT. Ctree is also
useable as a server but that version seems not to give trees as output.
Ward Wheeler and David Gladstein (wheeler@amnh.org) of the American Museum of Natural History, New York, have written
POY, a program that implements David Sankoff's method of
searching for the tree that minimizes a parsimony criterion that includes
penalties for gaps, accomplishing both searching for phylogenies and
alignments. POY has algorithmic improvements by Wheeler and Gladstein that
speed up the algorithm. (Their program MALIGN is the
only other program carrying out the full Sankoff proposal).
The method is described in a paper: Wheeler, W. 1996.
Optimization alignment: the end of multiple sequence alignment in
phylogenetics? Cladistics
12:1-9. POY is available in C source or in executables for
Linux, HPUX, SGI, Sun, and DOS. It is distributed by ftp from
ftp.amnh.org in directory pub/people/wheeler/poy.
Russell Doolittle (rdoolittle@ucsd.edu) and Dafei Feng, of the Department of Chemistry and Biochemistry of the University of California at San Diego, released ALIGN in 1990. A version for Macintoshes was coded by Peter Markeiwicz. ALIGN implements the "progressive alignment" strategy described in their paper: Feng, D.-F. and R. F. Doolittle. 1987. Progressive sequence aligment as a prerequisite to correct phylogenetic trees. Journal of Molecular Evolution 25: 351-360. This is also the basis for the Clustal family of programs as well as the Pileup program in the GCG package. The ALIGN program can align as well as print out a tree (which does not have branch lengths). It uses Doolittle's own formats, and so three other programs are included with ALIGN to convert formats. The programs are distributed by ftp from the EBI ftp software server at ftp.ebi.ac.uk in directory pub/software/mac as file align.hqx.
Pietro Lio, of the Department of Genetics,
University of Cambridge (P.Lio@gen.cam.ac.uk), has written
PASSML and PASSML_TM,
which use likelihood methods with Hidden Markov models to infer
phylogeny and also secondary structure from protein data. PASSML is for
general proteins and PASSML_TM is for membrane proteins.
The methods used are described in the papers: Goldman, N., J. L. Thorne,
and D. T. Jones. 1998. Assessing the impact of secondary structure and
solvent accessibility on protein evolution. Genetics 149:
445-458,
PASSML is described in the paper: Lio, P., N. Goldman, J. L. Thorne
and D. T. Jones. 1998. PASSML: combining evolutionary inference and protein
secondary structure prediction. Bioinformatics 14: 726-733,
and PASSML_TM is described in the paper:
Lio, P. and N. Goldman. 1999 Using protein structural information in
evolutionary inference: transmembrane proteins. Molecular Biology and
Evolution 16: 1696-1710.
The programs are
available as ANSI C for Unix workstations; PASSML has also been successfully
compiled under a port of gcc to a PC. The source code is available via the
web page at
http://ng-dec1.gen.cam.ac.uk/hmm/Passml.html.
Rod Page (dpage@udcf.gla.ac.uk) of the
Division of Environmental and Evolutionary Biology of the University
of Glasgow has written COMPONENT version 2.0, a program for
Windows systems for
comparing cladograms for use in phylogeny and biogeography studies. It has
many tree comparison and consensus methods, and far more features for
biogeographic studies (such as comparing species and area cladograms) than any
other package. It also can generate random trees. It runs under Windows 3.0
or higher. Its cost is 40 pounds U.K., and it can be ordered from
Anna Hutson at the
Department of Botany, Natural History Museum, Cromwell Road, London, SW7 5BD, U.K. (annah@nhm.ac.uk) (fax (0171) 938 9260).
Details on how to order may be found at the
order form
web page at http://taxonomy.zoology.gla.ac.uk/rod/order.html.
There is a review of the
program in Cladistics 9: 351-353 (1993). COMPONENT has a
web site at
http://taxonomy.zoology.gla.ac.uk/rod/cpw.html. The
documentation is available in Adobe Acrobat at
http://taxonomy.zoology.gla.ac.uk/rod/cplite/Manual.html.
A very early development Macintosh version ("COMPONENT Lite") is available free
from the
COMPONENT Lite web site
at http://taxonomy.zoology.gla.ac.uk/rod/cplite/guide.html
(though Rod says "be prepared for bugs").
Rod Page(dpage@udcf.gla.ac.uk), of the Division of Environmental and
Evolutionary Biology of the University of Glasgow has written TREEMAP, a free, experimental program for comparing host and parasite
phylogenies. It allows you to interactively compare host and parasite
trees, construct reconstructions of the history of the association, and
perform some simple randomisation tests of hypotheses of cospeciation.
The program is available as an executable for Macintoshes or an executable
for Windows PCs (the two versions are essentially identical).
They can be downloaded from
its WWW site: http://taxonomy.zoology.gla.ac.uk/rod/treemap.html.
The site also has an online manual, or you can download the documentation
as a Postscript file.
For a description of the method used by TreeMap, see Page, R.D.M. 1994.
Parallel Phylogenies: Reconstructing the history of host-parasite
assemblages. Cladistics 10: 155-173.