kr.angelov
|
1f77afcfce
|
the statistical parser now uses a baseline lexical estimation of the beam size
|
2013-02-12 09:41:32 +00:00 |
|
kr.angelov
|
a6b35a9053
|
the class PgfConcr from the Python binding now has a property name which returns the name of the concrete syntax
|
2013-02-11 15:51:26 +00:00 |
|
kr.angelov
|
0b7b939aca
|
refactoring: now all named objects in the C runtime have an explicit name field
|
2013-02-11 14:10:54 +00:00 |
|
kr.angelov
|
56c8f91d19
|
remove the pgf2yaml tool which was both broken and redundant. The declarations for generic programming from data.c are removed as well
|
2013-02-11 13:51:12 +00:00 |
|
kr.angelov
|
ff25ba8f90
|
the grammar reader in the C runtime is completely rewritten and it doesn't use the generic programming API
|
2013-02-11 10:16:58 +00:00 |
|
kr.angelov
|
be405532e6
|
a bunch of new words and fixes in DictEng and DictEngBul
|
2013-02-10 22:27:18 +00:00 |
|
kr.angelov
|
e9b5557c6c
|
This patch removes Gregoire's parse_tokens function in the python binding and adds another implementation which builds on the existing API for lexers in the C runtime. Now it is possible to write incremental Lexers in Python
|
2013-02-01 09:29:43 +00:00 |
|
kr.angelov
|
eca4a28563
|
implement gu_exn_caught in gu/exn.c. It was missing
|
2013-02-01 09:26:30 +00:00 |
|
kr.angelov
|
f4c56b7152
|
in NumeralAmh: UTF8 -> utf8. The former is not recognized on Windows
|
2013-01-31 15:08:13 +00:00 |
|
kr.angelov
|
7e5ad6eea6
|
fix the Windows link
|
2013-01-31 15:06:42 +00:00 |
|
kr.angelov
|
65bf4d9b9b
|
added a link to the Windows binary from the download page
|
2013-01-31 15:03:35 +00:00 |
|
kr.angelov
|
a4b0709923
|
a few fixes in DictEng
|
2013-01-30 09:58:39 +00:00 |
|
kr.angelov
|
6fce26c9dd
|
more words in DictEngBul.gf
|
2013-01-30 09:57:39 +00:00 |
|
kr.angelov
|
87545f3f83
|
bugfix in the reference counting for Python
|
2013-01-29 09:41:12 +00:00 |
|
kr.angelov
|
d4717d533a
|
the Python binding is in pure C again
|
2013-01-29 09:20:32 +00:00 |
|
kr.angelov
|
66282bfcb7
|
added an API for composing and decomposing abstract trees from Python
|
2013-01-29 09:07:41 +00:00 |
|
kr.angelov
|
1723d8637c
|
fixed typos in the python binding: in a few places pgf_ExprType was used instead of pgf_ExprIterType
|
2013-01-29 09:06:23 +00:00 |
|
kr.angelov
|
7b73100e01
|
switch from CP1251 to UTF8 in DictBul
|
2013-01-17 15:52:04 +00:00 |
|
kr.angelov
|
c222f86a2a
|
a few more words in DictEngBul.gf
|
2013-01-17 08:23:13 +00:00 |
|
kr.angelov
|
a7a7a21722
|
a few fixes in DictEngBul
|
2013-01-15 13:24:04 +00:00 |
|
kr.angelov
|
19288d0dda
|
about 3000 new words in DictEngBul.gf. The words are imported from the Universal WordNet but are not manually checked yet.
|
2013-01-15 11:17:59 +00:00 |
|
kr.angelov
|
ccc3d6be0d
|
fix warnings in pgf-parse.c
|
2013-01-08 12:53:49 +00:00 |
|
kr.angelov
|
79bf7056f2
|
now the Python binding has an alternative representation for abstract trees which is composed of Python objects. The new representation is not integrated with the core runtime yet
|
2013-01-07 15:11:12 +00:00 |
|
kr.angelov
|
3be31c62e9
|
a new reasoner in the C runtime. It supports tabling which makes it decideable for propositional logic. dependent types and high-order types are not supported yet. The generation is still in decreasing probability order
|
2013-01-07 12:50:32 +00:00 |
|
kr.angelov
|
0be179d7ff
|
bugfix in the strings library from the C runtime
|
2012-12-27 21:18:46 +00:00 |
|
kr.angelov
|
bb077b8330
|
bugfix: the linearizer should not generate extra space at the end of the sentence
|
2012-12-19 11:18:34 +00:00 |
|
kr.angelov
|
f7eaa8a89a
|
bugfix for linearization of metavariables at the root of a tree
|
2012-12-19 10:03:05 +00:00 |
|
kr.angelov
|
6201640d7b
|
rename linearize.{h/c} to linearizer.{h/c} which follows the convention used in parser.c and reasoner.c
|
2012-12-19 09:17:24 +00:00 |
|
kr.angelov
|
5c9ee467a9
|
a major reimplementation of the linearizer in the C runtime
|
2012-12-19 09:07:05 +00:00 |
|
kr.angelov
|
008c18a8a7
|
fixed accidental bug in pgf-parse.c
|
2012-12-18 15:42:04 +00:00 |
|
kr.angelov
|
dc809da91f
|
the C runtime now can read abstract expressions with literals and meta variables
|
2012-12-18 12:29:30 +00:00 |
|
kr.angelov
|
51d301d83c
|
updated statistical model
|
2012-12-17 10:31:04 +00:00 |
|
kr.angelov
|
a3f28fb521
|
some fixes in DictEng
|
2012-12-17 10:24:46 +00:00 |
|
kr.angelov
|
32905c8363
|
debugging infrastructure in the reasoner
|
2012-12-14 21:25:00 +00:00 |
|
kr.angelov
|
5cec2d5a50
|
bugfix for the reasoner in the C runtime
|
2012-12-14 21:24:17 +00:00 |
|
kr.angelov
|
b367dfd80f
|
a bit more flexible API for parsing in Python
|
2012-12-14 16:00:52 +00:00 |
|
kr.angelov
|
8aefd1e072
|
The first prototype for exhaustive generation in the C runtime. The trees are always listed in decreasing probability order. There is also an API for generation from Python
|
2012-12-14 15:32:49 +00:00 |
|
kr.angelov
|
e1bab39458
|
bugfix in the lexer from the C runtime. the input sentence doesn't have to terminate with whitespace
|
2012-12-13 16:45:44 +00:00 |
|
kr.angelov
|
6bc32db1c3
|
added simple error handling in the Python test
|
2012-12-13 16:44:39 +00:00 |
|
kr.angelov
|
81428c768c
|
added a simple test for the Python binding
|
2012-12-13 16:19:56 +00:00 |
|
kr.angelov
|
eebd9e92c9
|
a bugfix for building questions in the Bulgarian resource grammar
|
2012-12-13 15:52:55 +00:00 |
|
kr.angelov
|
cc7ea9260b
|
an initial API for parsing and linearization from Python
|
2012-12-13 15:39:07 +00:00 |
|
kr.angelov
|
2ba632dc9f
|
a top-level API for parsing in the C runtime
|
2012-12-13 14:44:33 +00:00 |
|
kr.angelov
|
60942c440a
|
bugfix: the outside probability of a PgfItemConts must always be initialized to zero
|
2012-12-13 11:11:45 +00:00 |
|
kr.angelov
|
fe51a7fb98
|
bugfix: pgf_read_expr no longer requires a semicolon at the end of an abstract expression
|
2012-12-13 11:09:26 +00:00 |
|
kr.angelov
|
162fd5e512
|
an initial Python binding to the C runtime
|
2012-12-12 11:29:39 +00:00 |
|
kr.angelov
|
1376df457d
|
started an official API to the C runtime
|
2012-12-12 11:25:58 +00:00 |
|
kr.angelov
|
3182e382dc
|
bugfix for robust parsing with multi-word units
|
2012-12-11 12:57:22 +00:00 |
|
kr.angelov
|
1863e4c3d6
|
added experimental script for chunking in the C runtime
|
2012-12-03 10:07:54 +00:00 |
|
kr.angelov
|
2da23e9872
|
added INSTALL file and updated README file for the C runtime
|
2012-12-03 09:09:08 +00:00 |
|
kr.angelov
|
818ea0d4d6
|
updated statistical model
|
2012-11-28 12:44:20 +00:00 |
|
kr.angelov
|
2db296db30
|
Pakistani is now properly capitalized in the English dictionary
|
2012-11-28 12:43:45 +00:00 |
|
kr.angelov
|
ee8f296089
|
translations for several words in the Bulgarian dictionary
|
2012-11-28 12:42:23 +00:00 |
|
kr.angelov
|
23873c8214
|
added familiar_A2 in DictEng, DictEngBul and DictEngGer
|
2012-11-28 12:40:20 +00:00 |
|
kr.angelov
|
0d35636348
|
bugfix for composite adjectivial phrases in Bulgarian
|
2012-11-28 12:37:24 +00:00 |
|
kr.angelov
|
6542f6ffdd
|
a bunch of additions in the parsing grammars and dictionaries plus an updated statistical model
|
2012-11-26 16:43:09 +00:00 |
|
kr.angelov
|
f8c302f9ef
|
remove the duplicated definition of PgfProductionIdx in parser.c
|
2012-11-19 14:16:31 +00:00 |
|
kr.angelov
|
ac2ea8a579
|
fix in ParseEngAbs3.probs
|
2012-11-19 14:14:08 +00:00 |
|
kr.angelov
|
6b0020f834
|
a statistical model for robustness in lib/src/english
|
2012-11-19 12:24:11 +00:00 |
|
kr.angelov
|
17ae9548d9
|
updated ParseEngAbs.probs with the latest statistics
|
2012-11-19 12:15:25 +00:00 |
|
kr.angelov
|
71b7c09ffe
|
bugfix for the building of bottom-up filter in the C runtime
|
2012-11-16 13:27:15 +00:00 |
|
kr.angelov
|
ba57ad3367
|
a couple of fixes and new words in DictEng and DictEngBul
|
2012-11-16 09:56:20 +00:00 |
|
kr.angelov
|
a3ba1991f4
|
revised heuristic in the statistical parser
|
2012-11-14 12:34:22 +00:00 |
|
kr.angelov
|
70c68f0527
|
bugfix in the statistical parser
|
2012-11-13 09:48:23 +00:00 |
|
kr.angelov
|
08ee662944
|
two simple heuristics which speed up the statistical parser more than seven times.
|
2012-11-12 22:17:40 +00:00 |
|
kr.angelov
|
68170d5b08
|
a simple refactoring in the statistical parser
|
2012-11-12 21:48:22 +00:00 |
|
kr.angelov
|
a2771552d6
|
more counters in the profiler for the statistical parser
|
2012-11-12 15:36:21 +00:00 |
|
kr.angelov
|
46de62c452
|
now we store the state instead of the offset for every continuation in the chart for the statistical parser
|
2012-11-12 14:04:52 +00:00 |
|
kr.angelov
|
9967c3ad04
|
in the statistical parser: move the outside probability from the parse items to their continuation. this makes the value slot shared between many items
|
2012-11-12 13:43:43 +00:00 |
|
kr.angelov
|
9d23093492
|
small refactoring in the C runtime
|
2012-11-12 13:05:35 +00:00 |
|
kr.angelov
|
a50c7c24b8
|
use size_t consistently as the type for constituent indices in the C runtime
|
2012-11-12 12:51:27 +00:00 |
|
kr.angelov
|
1e531e8237
|
implemented gu_map_count in runtime/c/gu/map.c
|
2012-11-12 12:42:19 +00:00 |
|
kr.angelov
|
52255664be
|
use prob_t instead of float in a few places
|
2012-10-29 08:52:56 +00:00 |
|
kr.angelov
|
0ad2405d69
|
forgot to add one #ifdef
|
2012-10-25 18:37:22 +00:00 |
|
kr.angelov
|
9721833680
|
a major refactoring in the robust parser: bottom-up filtering and garbage collection for the chart
|
2012-10-25 14:42:53 +00:00 |
|
kr.angelov
|
28b58b6267
|
add teyjus/simulator/builtins/builtins.h
|
2012-10-11 11:10:17 +00:00 |
|
kr.angelov
|
f0583bfd93
|
added the forgoten libteyjus.pc.in file in the C runtime
|
2012-10-11 04:22:38 +00:00 |
|
kr.angelov
|
bd08d98c7d
|
move examples/PennTreebank to /treebanks/PennTreebank
|
2012-10-01 08:52:54 +00:00 |
|
kr.angelov
|
00e85e55f8
|
Added as_Subj and UttAdV in the parsing grammars. Replaced plus_Prep with plus_Conj
|
2012-10-01 08:47:52 +00:00 |
|
kr.angelov
|
953633240e
|
typechecking and better error reporting in the training script for PennTreebank
|
2012-10-01 08:45:46 +00:00 |
|
kr.angelov
|
475109a40f
|
added the GF version of Talbanken which was imported by Malin
|
2012-10-01 08:33:43 +00:00 |
|
kr.angelov
|
6b7f1d2c6c
|
added AdvVPSlash and AdVVPSlash to VerbGer and an extended version of PPartNP which uses VPSlash in ParseEngGer. I guess the definitions so they might not be quite correct
|
2012-09-27 11:44:25 +00:00 |
|
kr.angelov
|
3845564625
|
added ParseEngGer.gf
|
2012-09-27 09:54:13 +00:00 |
|
kr.angelov
|
25e9f28fa4
|
add ApposNP and UncNeg to the Bulgarian parsing grammar
|
2012-09-27 09:31:37 +00:00 |
|
kr.angelov
|
933bdda844
|
use CNNumNP in the parsing grammars
|
2012-09-27 09:31:04 +00:00 |
|
kr.angelov
|
73823dbadc
|
remove no_RP from the parsing grammars and use EmptyRelSlash instead
|
2012-09-27 09:29:59 +00:00 |
|
kr.angelov
|
6084647328
|
added EmptyRelSlash in ExtraBul and ExtraGer. For Bulgarian and German the function simply inserts the default relative pronoun
|
2012-09-27 09:28:31 +00:00 |
|
kr.angelov
|
1b571d69ff
|
added according_to_Prep and ofter_AdA in DictEng, DictEngBul and DictEngGer
|
2012-09-27 09:11:04 +00:00 |
|
kr.angelov
|
e9800fa3eb
|
a few more words in DictEngBul
|
2012-09-27 09:09:49 +00:00 |
|
kr.angelov
|
dfc474580d
|
now in the parsing grammar ComplVV gets as additional arguments the polarity and the anteriority
|
2012-09-27 09:05:47 +00:00 |
|
kr.angelov
|
3f334fe321
|
an optimization in the German grammar for the dative/genitive variants
|
2012-09-26 11:11:42 +00:00 |
|
kr.angelov
|
e95e500b33
|
a bit of reordering in DictEngGer.gf
|
2012-09-26 09:16:17 +00:00 |
|
kr.angelov
|
b3f5835f8a
|
fixes in DictEngGer.gf
|
2012-09-26 08:52:18 +00:00 |
|
kr.angelov
|
fc89eaacca
|
260 new words in DictEngGer which are taken from the lexicon for patents
|
2012-09-26 08:26:39 +00:00 |
|
kr.angelov
|
8ba5e3fd64
|
fixes in the German parsing grammar and cleanup in DictEngGer.gf
|
2012-09-25 20:12:38 +00:00 |
|
kr.angelov
|
249d6cc2f8
|
fixes in the Bulgarian resource grammar. extensions in DictEng and DictEngBul
|
2012-09-24 09:41:14 +00:00 |
|
kr.angelov
|
18fe8af964
|
now the meta probability for a category is explicitly specified in the statistical model instead of computed internally. this avoids rounding errors while computing the sum of a large number of small values.
|
2012-09-24 09:37:21 +00:00 |
|
kr.angelov
|
bb15542a85
|
in the robust parser we don't have to care about trees which yeld empty strings. this makes the parser a lot faster
|
2012-09-24 09:30:20 +00:00 |
|
kr.angelov
|
f75d1374ff
|
the Haskell runtime now exports 'functionsByCat' which returns the list of all functions for a given category
|
2012-09-18 09:48:21 +00:00 |
|
kr.angelov
|
e98d62c42a
|
a few changes in DictEng
|
2012-09-18 09:25:22 +00:00 |
|