kr.angelov
|
b47dfd9dbb
|
forgot to add reasoner.h
|
2013-06-26 09:09:54 +00:00 |
|
kr.angelov
|
67872578c9
|
forgot to add jit.h
|
2013-06-26 09:08:47 +00:00 |
|
kr.angelov
|
c873531172
|
an optimization in the jitter for generating more compact code
|
2013-06-26 09:03:51 +00:00 |
|
kr.angelov
|
dba75911b0
|
patch for adjustable heuristics from Python
|
2013-06-26 07:36:03 +00:00 |
|
kr.angelov
|
966d3aee3f
|
compatibility issue for MacOS X
|
2013-06-26 07:03:32 +00:00 |
|
kr.angelov
|
38b3dfcad6
|
fix for x86_64
|
2013-06-26 06:43:33 +00:00 |
|
kr.angelov
|
16584d4368
|
Now there is a just-in-time compiler which generates native code for proof search. This is already used by the exhaustive generator. The time to generate 10000 abstract trees with ParseEng went down from 4.43 sec to 0.29 sec.
|
2013-06-25 19:22:42 +00:00 |
|
kr.angelov
|
10eb9dedb6
|
bugfix for the linearizer in the C runtime
|
2013-06-24 07:56:42 +00:00 |
|
kr.angelov
|
c210da79a9
|
extensions in DictEngBul.gf
|
2013-06-22 15:41:52 +00:00 |
|
kr.angelov
|
aacc15b58f
|
bugfix for the word completion in the C runtime
|
2013-06-22 15:39:47 +00:00 |
|
kr.angelov
|
72cd14a5ae
|
add x86_64 support to GNU lightning
|
2013-06-20 08:27:04 +00:00 |
|
kr.angelov
|
e720d47700
|
fíx in the Python binding for compatibility with Python<2.7
|
2013-06-19 19:31:55 +00:00 |
|
kr.angelov
|
eece31c1ab
|
fix an issue in the Python binding related to the 32 vs 64 bit compatibility
|
2013-06-19 18:47:52 +00:00 |
|
kr.angelov
|
ffa6cbd03c
|
fix for a couple of warnings that are generated when GNU lightning is used
|
2013-06-17 07:32:41 +00:00 |
|
kr.angelov
|
6c4f52faeb
|
add the source code for GNU lightning in the source directory for the C runtime
|
2013-06-17 07:26:00 +00:00 |
|
kr.angelov
|
c3b344084f
|
bugfix in the python binding
|
2013-06-14 07:02:53 +00:00 |
|
kr.angelov
|
1b791158af
|
now the call Expr.unpack("? e1 e2") in Python returns a pair with None as the first element and a the list [e1,e2] as the second. This makes it easier to decompose partial abstract trees
|
2013-05-31 09:24:15 +00:00 |
|
kr.angelov
|
3566143f37
|
improved error message in the Python binding
|
2013-05-31 09:13:16 +00:00 |
|
kr.angelov
|
935ae49376
|
bugfix for the grammar printer in the C runtime
|
2013-05-30 20:20:02 +00:00 |
|
kr.angelov
|
0330a2e5e8
|
the Bulgarian phrasebook works again
|
2013-05-30 13:16:35 +00:00 |
|
kr.angelov
|
d66dfe13c2
|
a simple refactoring in the Python runtime
|
2013-05-29 11:02:18 +00:00 |
|
kr.angelov
|
fac39a78fe
|
readPGF in the Python runtime now throws "No such file or directory" exception if the grammar is missing
|
2013-05-29 10:49:56 +00:00 |
|
kr.angelov
|
e969aa69ff
|
added a test class for the Java API plus a small refinement in the implementation for the binding
|
2013-05-28 13:32:32 +00:00 |
|
kr.angelov
|
bd859fcf28
|
an initial skeleton for building a Java binding to the C runtime
|
2013-05-28 12:59:19 +00:00 |
|
kr.angelov
|
af8cec11f9
|
substantive and relative nouns in the Bulgarian library
|
2013-05-23 10:15:10 +00:00 |
|
kr.angelov
|
0a66a17e98
|
a bunch of changes in DictEng and DictEngBul
|
2013-05-23 09:29:59 +00:00 |
|
kr.angelov
|
9a6b504407
|
remove DiffBul.gf which is now obsolete
|
2013-05-23 08:59:36 +00:00 |
|
kr.angelov
|
78aab96369
|
fix the encoding problem with unicode literals in the Python binding
|
2013-05-21 10:53:20 +00:00 |
|
kr.angelov
|
f9f0fdcdf8
|
bugfix for bracketedLinearize which was causing crash if the tree cannot be linearized
|
2013-05-07 08:35:33 +00:00 |
|
kr.angelov
|
2eb37f6407
|
bug fix in the management of memory pools in the statistical parser
|
2013-05-07 08:30:32 +00:00 |
|
kr.angelov
|
561e478ed4
|
the statistical parser is now using two memory pools: one for parsing and one for the output trees. This means that the memory for parsing can be released as soon as the needed abstract trees are retrieved, while the trees themselves are retained in the separate output pool
|
2013-05-06 15:28:04 +00:00 |
|
kr.angelov
|
307e0854ed
|
fix the leftcorner filtering after the addition of word completion
|
2013-05-05 10:30:06 +00:00 |
|
kr.angelov
|
be8d72d64c
|
bugfix in the C runtime which was causing an infinite loop while linearizing partial trees
|
2013-05-04 13:32:57 +00:00 |
|
kr.angelov
|
b7ada4d269
|
remove /lib/src/english/ParseEngAbs3.probs since now it is moved to /treebanks/PennTreebank/ParseEngAbs3.probs
|
2013-05-03 21:25:31 +00:00 |
|
kr.angelov
|
9cdd96363a
|
word completion in the C runtime. The runtime/python/test.py example is now using readline with word completion
|
2013-05-01 06:09:55 +00:00 |
|
kr.angelov
|
6cc44193b8
|
finally the statistical parser is able to return all possible abstract trees
|
2013-04-26 20:44:01 +00:00 |
|
kr.angelov
|
f4cf8deab7
|
a trivial refactoring of the reasoner in the C runtime
|
2013-04-23 06:40:14 +00:00 |
|
kr.angelov
|
5aee2c4473
|
bug fix in pgf-translate which was hiding that there are more than one trees per sentence
|
2013-04-22 13:02:43 +00:00 |
|
kr.angelov
|
e76873c11f
|
a bit more informative error message in GrammarToPGF
|
2013-04-22 12:14:39 +00:00 |
|
kr.angelov
|
ffd64cc02a
|
reverse the direction of the arcs in the dependency trees
|
2013-04-21 19:20:08 +00:00 |
|
kr.angelov
|
23bf85e023
|
the option -old for the vp command is now redundant
|
2013-04-19 11:15:18 +00:00 |
|
kr.angelov
|
d6d4ae3a6b
|
remove the dead code left behind by Peter Ljunglöf in VisualizeTree
|
2013-04-19 11:13:07 +00:00 |
|
kr.angelov
|
b4374a6a52
|
fix the command options for the vd command in the shell
|
2013-04-19 11:11:57 +00:00 |
|
kr.angelov
|
15fd8b15ab
|
the C runtime and the Python binding now have an API for parser evaluation. The API computes PARSEVAL and Exact Match for a given tree. As a side effect the abstract trees in Python are now compared for equality by value and not by reference
|
2013-04-19 10:57:46 +00:00 |
|
kr.angelov
|
2a0c69a412
|
added API for computing bracketed strings from Python and C
|
2013-04-18 13:37:09 +00:00 |
|
kr.angelov
|
cb7025dc11
|
added a malt_tab format to the vd command in the GF shell
|
2013-04-16 18:22:37 +00:00 |
|
kr.angelov
|
d5666aebd0
|
the generation of dependency trees in the Haskell runtime is now finally working with bracketed strings. This also fixes some errors in the old implementation
|
2013-04-16 13:10:48 +00:00 |
|
kr.angelov
|
44828765c3
|
the compiler now sorts the list of functions per category in probability order. this ensures probability order search in the C runtime
|
2013-04-15 19:58:57 +00:00 |
|
kr.angelov
|
b6bbe96503
|
now the web service to the robust parser can to translations also
|
2013-04-05 12:22:52 +00:00 |
|
kr.angelov
|
cf0da12b8a
|
a bugfix which was causing an infinite loop in the C linearizer for some sentences
|
2013-04-05 09:11:24 +00:00 |
|
kr.angelov
|
b850ea2b9b
|
a very simple linearization for partial abstract trees in the C runtime
|
2013-04-05 08:42:56 +00:00 |
|
kr.angelov
|
f11abff7c6
|
added simple script for estimating the coverage on the PennTreebank
|
2013-03-28 09:15:38 +00:00 |
|
kr.angelov
|
ad4c97fdf7
|
added a few more multiword expressions in DictEng and a few words in the abstract syntax are not tagged with their senses. There is a new statistical model too
|
2013-03-27 20:46:42 +00:00 |
|
kr.angelov
|
be922d09a1
|
added the file treebanks/PennTreebank/ParseEngAbs3.probs which is used by the statistical parser for robust chunking
|
2013-03-25 10:28:53 +00:00 |
|
kr.angelov
|
72556ad1ae
|
a long list of prepositions from Wikipedia is now imported in DictEng in addition there are a number of small other changes in the dictionary. The statistical model is updated and is now moved to treebanks/PennTreebank/ParseEngAbs.probs
|
2013-03-25 10:24:24 +00:00 |
|
kr.angelov
|
8b40d4974b
|
added configuration file which defines the heads for all syntactic functions in ParseEng
|
2013-03-21 13:39:24 +00:00 |
|
kr.angelov
|
650e1cfa43
|
the calculation of lexical_prob in the statistical parser doesn't work properly. It should be fixed but for now I just disabled the optimization
|
2013-03-20 12:28:52 +00:00 |
|
kr.angelov
|
fec34e7622
|
replace #if with #ifdef when checking for the optional bottom up filtering in the C runtime
|
2013-03-20 10:47:47 +00:00 |
|
kr.angelov
|
466813f1e8
|
fix in ParseHin which made it impossible to load the grammar with the C runtime
|
2013-03-20 10:34:37 +00:00 |
|
kr.angelov
|
1ddcfc219e
|
the bottom up filtering in the C runtime is temporary disabled. It takes too much memory and even makes it impossible to load the Finnish and the German parsing grammars.
|
2013-03-19 10:59:44 +00:00 |
|
kr.angelov
|
8041999405
|
the ParseFin grammar now excludes ComplVV from VerbFin since this function has a more general type in the parsing grammar
|
2013-03-19 10:49:13 +00:00 |
|
kr.angelov
|
c775d0c5c5
|
filterout all adjectives and adverbs which could be derived morphologically
|
2013-03-18 17:31:20 +00:00 |
|
kr.angelov
|
34fddf669f
|
some of the newly added nouns in DictEng were actually variations of already existing lexical entries. Those are removed now.
|
2013-03-15 23:23:06 +00:00 |
|
kr.angelov
|
e5913189db
|
massive extensions in DictEng and DictEngBul. This includes all new nouns imported from WordNet by Shafqat, phrasal verbs that I collected from internet and the PennTreebank, plus various other small additions.
|
2013-03-15 20:18:22 +00:00 |
|
kr.angelov
|
cb37254882
|
bug fix in the linearizer in the C runtime
|
2013-03-14 12:31:49 +00:00 |
|
kr.angelov
|
f1a42ad78e
|
update the pgf-service tool from the C runtime after the changes in the API
|
2013-03-14 10:37:01 +00:00 |
|
kr.angelov
|
2893397fbb
|
bugfix in the statistical parser
|
2013-03-11 14:47:43 +00:00 |
|
kr.angelov
|
d924b70888
|
a bunch of changes in DictEng and DictEngBul plus an updated statistical model
|
2013-02-27 09:21:29 +00:00 |
|
kr.angelov
|
f001d40ae3
|
added gu_buf_flush in seq.c which removes all elements from a buffer
|
2013-02-26 09:48:09 +00:00 |
|
kr.angelov
|
1a0f85d297
|
fixes and extensions in DictEng and DictBul
|
2013-02-22 09:53:54 +00:00 |
|
kr.angelov
|
5a54596fe8
|
the parser in the C runtime should not crash if the start category is not defined
|
2013-02-19 12:08:48 +00:00 |
|
kr.angelov
|
f86dcb6572
|
bugfix in the grammar reader in the C runtime
|
2013-02-19 12:04:10 +00:00 |
|
kr.angelov
|
ffb17bd26a
|
bugfix in the linearizer for the C runtime
|
2013-02-13 15:39:01 +00:00 |
|
kr.angelov
|
55203110bb
|
now the beam size for the statistical parser can be configured by using the flag beam_size in the top-level concrete module
|
2013-02-12 10:53:13 +00:00 |
|
kr.angelov
|
1f77afcfce
|
the statistical parser now uses a baseline lexical estimation of the beam size
|
2013-02-12 09:41:32 +00:00 |
|
kr.angelov
|
a6b35a9053
|
the class PgfConcr from the Python binding now has a property name which returns the name of the concrete syntax
|
2013-02-11 15:51:26 +00:00 |
|
kr.angelov
|
0b7b939aca
|
refactoring: now all named objects in the C runtime have an explicit name field
|
2013-02-11 14:10:54 +00:00 |
|
kr.angelov
|
56c8f91d19
|
remove the pgf2yaml tool which was both broken and redundant. The declarations for generic programming from data.c are removed as well
|
2013-02-11 13:51:12 +00:00 |
|
kr.angelov
|
ff25ba8f90
|
the grammar reader in the C runtime is completely rewritten and it doesn't use the generic programming API
|
2013-02-11 10:16:58 +00:00 |
|
kr.angelov
|
be405532e6
|
a bunch of new words and fixes in DictEng and DictEngBul
|
2013-02-10 22:27:18 +00:00 |
|
kr.angelov
|
e9b5557c6c
|
This patch removes Gregoire's parse_tokens function in the python binding and adds another implementation which builds on the existing API for lexers in the C runtime. Now it is possible to write incremental Lexers in Python
|
2013-02-01 09:29:43 +00:00 |
|
kr.angelov
|
eca4a28563
|
implement gu_exn_caught in gu/exn.c. It was missing
|
2013-02-01 09:26:30 +00:00 |
|
kr.angelov
|
f4c56b7152
|
in NumeralAmh: UTF8 -> utf8. The former is not recognized on Windows
|
2013-01-31 15:08:13 +00:00 |
|
kr.angelov
|
7e5ad6eea6
|
fix the Windows link
|
2013-01-31 15:06:42 +00:00 |
|
kr.angelov
|
65bf4d9b9b
|
added a link to the Windows binary from the download page
|
2013-01-31 15:03:35 +00:00 |
|
kr.angelov
|
a4b0709923
|
a few fixes in DictEng
|
2013-01-30 09:58:39 +00:00 |
|
kr.angelov
|
6fce26c9dd
|
more words in DictEngBul.gf
|
2013-01-30 09:57:39 +00:00 |
|
kr.angelov
|
87545f3f83
|
bugfix in the reference counting for Python
|
2013-01-29 09:41:12 +00:00 |
|
kr.angelov
|
d4717d533a
|
the Python binding is in pure C again
|
2013-01-29 09:20:32 +00:00 |
|
kr.angelov
|
66282bfcb7
|
added an API for composing and decomposing abstract trees from Python
|
2013-01-29 09:07:41 +00:00 |
|
kr.angelov
|
1723d8637c
|
fixed typos in the python binding: in a few places pgf_ExprType was used instead of pgf_ExprIterType
|
2013-01-29 09:06:23 +00:00 |
|
kr.angelov
|
7b73100e01
|
switch from CP1251 to UTF8 in DictBul
|
2013-01-17 15:52:04 +00:00 |
|
kr.angelov
|
c222f86a2a
|
a few more words in DictEngBul.gf
|
2013-01-17 08:23:13 +00:00 |
|
kr.angelov
|
a7a7a21722
|
a few fixes in DictEngBul
|
2013-01-15 13:24:04 +00:00 |
|
kr.angelov
|
19288d0dda
|
about 3000 new words in DictEngBul.gf. The words are imported from the Universal WordNet but are not manually checked yet.
|
2013-01-15 11:17:59 +00:00 |
|
kr.angelov
|
ccc3d6be0d
|
fix warnings in pgf-parse.c
|
2013-01-08 12:53:49 +00:00 |
|
kr.angelov
|
79bf7056f2
|
now the Python binding has an alternative representation for abstract trees which is composed of Python objects. The new representation is not integrated with the core runtime yet
|
2013-01-07 15:11:12 +00:00 |
|
kr.angelov
|
3be31c62e9
|
a new reasoner in the C runtime. It supports tabling which makes it decideable for propositional logic. dependent types and high-order types are not supported yet. The generation is still in decreasing probability order
|
2013-01-07 12:50:32 +00:00 |
|
kr.angelov
|
0be179d7ff
|
bugfix in the strings library from the C runtime
|
2012-12-27 21:18:46 +00:00 |
|
kr.angelov
|
bb077b8330
|
bugfix: the linearizer should not generate extra space at the end of the sentence
|
2012-12-19 11:18:34 +00:00 |
|