kr.angelov
|
fac39a78fe
|
readPGF in the Python runtime now throws "No such file or directory" exception if the grammar is missing
|
2013-05-29 10:49:56 +00:00 |
|
kr.angelov
|
e969aa69ff
|
added a test class for the Java API plus a small refinement in the implementation for the binding
|
2013-05-28 13:32:32 +00:00 |
|
kr.angelov
|
bd859fcf28
|
an initial skeleton for building a Java binding to the C runtime
|
2013-05-28 12:59:19 +00:00 |
|
kr.angelov
|
af8cec11f9
|
substantive and relative nouns in the Bulgarian library
|
2013-05-23 10:15:10 +00:00 |
|
kr.angelov
|
0a66a17e98
|
a bunch of changes in DictEng and DictEngBul
|
2013-05-23 09:29:59 +00:00 |
|
kr.angelov
|
9a6b504407
|
remove DiffBul.gf which is now obsolete
|
2013-05-23 08:59:36 +00:00 |
|
kr.angelov
|
78aab96369
|
fix the encoding problem with unicode literals in the Python binding
|
2013-05-21 10:53:20 +00:00 |
|
kr.angelov
|
f9f0fdcdf8
|
bugfix for bracketedLinearize which was causing crash if the tree cannot be linearized
|
2013-05-07 08:35:33 +00:00 |
|
kr.angelov
|
2eb37f6407
|
bug fix in the management of memory pools in the statistical parser
|
2013-05-07 08:30:32 +00:00 |
|
kr.angelov
|
561e478ed4
|
the statistical parser is now using two memory pools: one for parsing and one for the output trees. This means that the memory for parsing can be released as soon as the needed abstract trees are retrieved, while the trees themselves are retained in the separate output pool
|
2013-05-06 15:28:04 +00:00 |
|
kr.angelov
|
307e0854ed
|
fix the leftcorner filtering after the addition of word completion
|
2013-05-05 10:30:06 +00:00 |
|
kr.angelov
|
be8d72d64c
|
bugfix in the C runtime which was causing an infinite loop while linearizing partial trees
|
2013-05-04 13:32:57 +00:00 |
|
kr.angelov
|
b7ada4d269
|
remove /lib/src/english/ParseEngAbs3.probs since now it is moved to /treebanks/PennTreebank/ParseEngAbs3.probs
|
2013-05-03 21:25:31 +00:00 |
|
kr.angelov
|
9cdd96363a
|
word completion in the C runtime. The runtime/python/test.py example is now using readline with word completion
|
2013-05-01 06:09:55 +00:00 |
|
kr.angelov
|
6cc44193b8
|
finally the statistical parser is able to return all possible abstract trees
|
2013-04-26 20:44:01 +00:00 |
|
kr.angelov
|
f4cf8deab7
|
a trivial refactoring of the reasoner in the C runtime
|
2013-04-23 06:40:14 +00:00 |
|
kr.angelov
|
5aee2c4473
|
bug fix in pgf-translate which was hiding that there are more than one trees per sentence
|
2013-04-22 13:02:43 +00:00 |
|
kr.angelov
|
e76873c11f
|
a bit more informative error message in GrammarToPGF
|
2013-04-22 12:14:39 +00:00 |
|
kr.angelov
|
ffd64cc02a
|
reverse the direction of the arcs in the dependency trees
|
2013-04-21 19:20:08 +00:00 |
|
kr.angelov
|
23bf85e023
|
the option -old for the vp command is now redundant
|
2013-04-19 11:15:18 +00:00 |
|
kr.angelov
|
d6d4ae3a6b
|
remove the dead code left behind by Peter Ljunglöf in VisualizeTree
|
2013-04-19 11:13:07 +00:00 |
|
kr.angelov
|
b4374a6a52
|
fix the command options for the vd command in the shell
|
2013-04-19 11:11:57 +00:00 |
|
kr.angelov
|
15fd8b15ab
|
the C runtime and the Python binding now have an API for parser evaluation. The API computes PARSEVAL and Exact Match for a given tree. As a side effect the abstract trees in Python are now compared for equality by value and not by reference
|
2013-04-19 10:57:46 +00:00 |
|
kr.angelov
|
2a0c69a412
|
added API for computing bracketed strings from Python and C
|
2013-04-18 13:37:09 +00:00 |
|
kr.angelov
|
cb7025dc11
|
added a malt_tab format to the vd command in the GF shell
|
2013-04-16 18:22:37 +00:00 |
|
kr.angelov
|
d5666aebd0
|
the generation of dependency trees in the Haskell runtime is now finally working with bracketed strings. This also fixes some errors in the old implementation
|
2013-04-16 13:10:48 +00:00 |
|
kr.angelov
|
44828765c3
|
the compiler now sorts the list of functions per category in probability order. this ensures probability order search in the C runtime
|
2013-04-15 19:58:57 +00:00 |
|
kr.angelov
|
b6bbe96503
|
now the web service to the robust parser can to translations also
|
2013-04-05 12:22:52 +00:00 |
|
kr.angelov
|
cf0da12b8a
|
a bugfix which was causing an infinite loop in the C linearizer for some sentences
|
2013-04-05 09:11:24 +00:00 |
|
kr.angelov
|
b850ea2b9b
|
a very simple linearization for partial abstract trees in the C runtime
|
2013-04-05 08:42:56 +00:00 |
|
kr.angelov
|
f11abff7c6
|
added simple script for estimating the coverage on the PennTreebank
|
2013-03-28 09:15:38 +00:00 |
|
kr.angelov
|
ad4c97fdf7
|
added a few more multiword expressions in DictEng and a few words in the abstract syntax are not tagged with their senses. There is a new statistical model too
|
2013-03-27 20:46:42 +00:00 |
|
kr.angelov
|
be922d09a1
|
added the file treebanks/PennTreebank/ParseEngAbs3.probs which is used by the statistical parser for robust chunking
|
2013-03-25 10:28:53 +00:00 |
|
kr.angelov
|
72556ad1ae
|
a long list of prepositions from Wikipedia is now imported in DictEng in addition there are a number of small other changes in the dictionary. The statistical model is updated and is now moved to treebanks/PennTreebank/ParseEngAbs.probs
|
2013-03-25 10:24:24 +00:00 |
|
kr.angelov
|
8b40d4974b
|
added configuration file which defines the heads for all syntactic functions in ParseEng
|
2013-03-21 13:39:24 +00:00 |
|
kr.angelov
|
650e1cfa43
|
the calculation of lexical_prob in the statistical parser doesn't work properly. It should be fixed but for now I just disabled the optimization
|
2013-03-20 12:28:52 +00:00 |
|
kr.angelov
|
fec34e7622
|
replace #if with #ifdef when checking for the optional bottom up filtering in the C runtime
|
2013-03-20 10:47:47 +00:00 |
|
kr.angelov
|
466813f1e8
|
fix in ParseHin which made it impossible to load the grammar with the C runtime
|
2013-03-20 10:34:37 +00:00 |
|
kr.angelov
|
1ddcfc219e
|
the bottom up filtering in the C runtime is temporary disabled. It takes too much memory and even makes it impossible to load the Finnish and the German parsing grammars.
|
2013-03-19 10:59:44 +00:00 |
|
kr.angelov
|
8041999405
|
the ParseFin grammar now excludes ComplVV from VerbFin since this function has a more general type in the parsing grammar
|
2013-03-19 10:49:13 +00:00 |
|
kr.angelov
|
c775d0c5c5
|
filterout all adjectives and adverbs which could be derived morphologically
|
2013-03-18 17:31:20 +00:00 |
|
kr.angelov
|
34fddf669f
|
some of the newly added nouns in DictEng were actually variations of already existing lexical entries. Those are removed now.
|
2013-03-15 23:23:06 +00:00 |
|
kr.angelov
|
e5913189db
|
massive extensions in DictEng and DictEngBul. This includes all new nouns imported from WordNet by Shafqat, phrasal verbs that I collected from internet and the PennTreebank, plus various other small additions.
|
2013-03-15 20:18:22 +00:00 |
|
kr.angelov
|
cb37254882
|
bug fix in the linearizer in the C runtime
|
2013-03-14 12:31:49 +00:00 |
|
kr.angelov
|
f1a42ad78e
|
update the pgf-service tool from the C runtime after the changes in the API
|
2013-03-14 10:37:01 +00:00 |
|
kr.angelov
|
2893397fbb
|
bugfix in the statistical parser
|
2013-03-11 14:47:43 +00:00 |
|
kr.angelov
|
d924b70888
|
a bunch of changes in DictEng and DictEngBul plus an updated statistical model
|
2013-02-27 09:21:29 +00:00 |
|
kr.angelov
|
f001d40ae3
|
added gu_buf_flush in seq.c which removes all elements from a buffer
|
2013-02-26 09:48:09 +00:00 |
|
kr.angelov
|
1a0f85d297
|
fixes and extensions in DictEng and DictBul
|
2013-02-22 09:53:54 +00:00 |
|
kr.angelov
|
5a54596fe8
|
the parser in the C runtime should not crash if the start category is not defined
|
2013-02-19 12:08:48 +00:00 |
|
kr.angelov
|
f86dcb6572
|
bugfix in the grammar reader in the C runtime
|
2013-02-19 12:04:10 +00:00 |
|
kr.angelov
|
ffb17bd26a
|
bugfix in the linearizer for the C runtime
|
2013-02-13 15:39:01 +00:00 |
|
kr.angelov
|
55203110bb
|
now the beam size for the statistical parser can be configured by using the flag beam_size in the top-level concrete module
|
2013-02-12 10:53:13 +00:00 |
|
kr.angelov
|
1f77afcfce
|
the statistical parser now uses a baseline lexical estimation of the beam size
|
2013-02-12 09:41:32 +00:00 |
|
kr.angelov
|
a6b35a9053
|
the class PgfConcr from the Python binding now has a property name which returns the name of the concrete syntax
|
2013-02-11 15:51:26 +00:00 |
|
kr.angelov
|
0b7b939aca
|
refactoring: now all named objects in the C runtime have an explicit name field
|
2013-02-11 14:10:54 +00:00 |
|
kr.angelov
|
56c8f91d19
|
remove the pgf2yaml tool which was both broken and redundant. The declarations for generic programming from data.c are removed as well
|
2013-02-11 13:51:12 +00:00 |
|
kr.angelov
|
ff25ba8f90
|
the grammar reader in the C runtime is completely rewritten and it doesn't use the generic programming API
|
2013-02-11 10:16:58 +00:00 |
|
kr.angelov
|
be405532e6
|
a bunch of new words and fixes in DictEng and DictEngBul
|
2013-02-10 22:27:18 +00:00 |
|
kr.angelov
|
e9b5557c6c
|
This patch removes Gregoire's parse_tokens function in the python binding and adds another implementation which builds on the existing API for lexers in the C runtime. Now it is possible to write incremental Lexers in Python
|
2013-02-01 09:29:43 +00:00 |
|
kr.angelov
|
eca4a28563
|
implement gu_exn_caught in gu/exn.c. It was missing
|
2013-02-01 09:26:30 +00:00 |
|
kr.angelov
|
f4c56b7152
|
in NumeralAmh: UTF8 -> utf8. The former is not recognized on Windows
|
2013-01-31 15:08:13 +00:00 |
|
kr.angelov
|
7e5ad6eea6
|
fix the Windows link
|
2013-01-31 15:06:42 +00:00 |
|
kr.angelov
|
65bf4d9b9b
|
added a link to the Windows binary from the download page
|
2013-01-31 15:03:35 +00:00 |
|
kr.angelov
|
a4b0709923
|
a few fixes in DictEng
|
2013-01-30 09:58:39 +00:00 |
|
kr.angelov
|
6fce26c9dd
|
more words in DictEngBul.gf
|
2013-01-30 09:57:39 +00:00 |
|
kr.angelov
|
87545f3f83
|
bugfix in the reference counting for Python
|
2013-01-29 09:41:12 +00:00 |
|
kr.angelov
|
d4717d533a
|
the Python binding is in pure C again
|
2013-01-29 09:20:32 +00:00 |
|
kr.angelov
|
66282bfcb7
|
added an API for composing and decomposing abstract trees from Python
|
2013-01-29 09:07:41 +00:00 |
|
kr.angelov
|
1723d8637c
|
fixed typos in the python binding: in a few places pgf_ExprType was used instead of pgf_ExprIterType
|
2013-01-29 09:06:23 +00:00 |
|
kr.angelov
|
7b73100e01
|
switch from CP1251 to UTF8 in DictBul
|
2013-01-17 15:52:04 +00:00 |
|
kr.angelov
|
c222f86a2a
|
a few more words in DictEngBul.gf
|
2013-01-17 08:23:13 +00:00 |
|
kr.angelov
|
a7a7a21722
|
a few fixes in DictEngBul
|
2013-01-15 13:24:04 +00:00 |
|
kr.angelov
|
19288d0dda
|
about 3000 new words in DictEngBul.gf. The words are imported from the Universal WordNet but are not manually checked yet.
|
2013-01-15 11:17:59 +00:00 |
|
kr.angelov
|
ccc3d6be0d
|
fix warnings in pgf-parse.c
|
2013-01-08 12:53:49 +00:00 |
|
kr.angelov
|
79bf7056f2
|
now the Python binding has an alternative representation for abstract trees which is composed of Python objects. The new representation is not integrated with the core runtime yet
|
2013-01-07 15:11:12 +00:00 |
|
kr.angelov
|
3be31c62e9
|
a new reasoner in the C runtime. It supports tabling which makes it decideable for propositional logic. dependent types and high-order types are not supported yet. The generation is still in decreasing probability order
|
2013-01-07 12:50:32 +00:00 |
|
kr.angelov
|
0be179d7ff
|
bugfix in the strings library from the C runtime
|
2012-12-27 21:18:46 +00:00 |
|
kr.angelov
|
bb077b8330
|
bugfix: the linearizer should not generate extra space at the end of the sentence
|
2012-12-19 11:18:34 +00:00 |
|
kr.angelov
|
f7eaa8a89a
|
bugfix for linearization of metavariables at the root of a tree
|
2012-12-19 10:03:05 +00:00 |
|
kr.angelov
|
6201640d7b
|
rename linearize.{h/c} to linearizer.{h/c} which follows the convention used in parser.c and reasoner.c
|
2012-12-19 09:17:24 +00:00 |
|
kr.angelov
|
5c9ee467a9
|
a major reimplementation of the linearizer in the C runtime
|
2012-12-19 09:07:05 +00:00 |
|
kr.angelov
|
008c18a8a7
|
fixed accidental bug in pgf-parse.c
|
2012-12-18 15:42:04 +00:00 |
|
kr.angelov
|
dc809da91f
|
the C runtime now can read abstract expressions with literals and meta variables
|
2012-12-18 12:29:30 +00:00 |
|
kr.angelov
|
51d301d83c
|
updated statistical model
|
2012-12-17 10:31:04 +00:00 |
|
kr.angelov
|
a3f28fb521
|
some fixes in DictEng
|
2012-12-17 10:24:46 +00:00 |
|
kr.angelov
|
32905c8363
|
debugging infrastructure in the reasoner
|
2012-12-14 21:25:00 +00:00 |
|
kr.angelov
|
5cec2d5a50
|
bugfix for the reasoner in the C runtime
|
2012-12-14 21:24:17 +00:00 |
|
kr.angelov
|
b367dfd80f
|
a bit more flexible API for parsing in Python
|
2012-12-14 16:00:52 +00:00 |
|
kr.angelov
|
8aefd1e072
|
The first prototype for exhaustive generation in the C runtime. The trees are always listed in decreasing probability order. There is also an API for generation from Python
|
2012-12-14 15:32:49 +00:00 |
|
kr.angelov
|
e1bab39458
|
bugfix in the lexer from the C runtime. the input sentence doesn't have to terminate with whitespace
|
2012-12-13 16:45:44 +00:00 |
|
kr.angelov
|
6bc32db1c3
|
added simple error handling in the Python test
|
2012-12-13 16:44:39 +00:00 |
|
kr.angelov
|
81428c768c
|
added a simple test for the Python binding
|
2012-12-13 16:19:56 +00:00 |
|
kr.angelov
|
eebd9e92c9
|
a bugfix for building questions in the Bulgarian resource grammar
|
2012-12-13 15:52:55 +00:00 |
|
kr.angelov
|
cc7ea9260b
|
an initial API for parsing and linearization from Python
|
2012-12-13 15:39:07 +00:00 |
|
kr.angelov
|
2ba632dc9f
|
a top-level API for parsing in the C runtime
|
2012-12-13 14:44:33 +00:00 |
|
kr.angelov
|
60942c440a
|
bugfix: the outside probability of a PgfItemConts must always be initialized to zero
|
2012-12-13 11:11:45 +00:00 |
|
kr.angelov
|
fe51a7fb98
|
bugfix: pgf_read_expr no longer requires a semicolon at the end of an abstract expression
|
2012-12-13 11:09:26 +00:00 |
|
kr.angelov
|
162fd5e512
|
an initial Python binding to the C runtime
|
2012-12-12 11:29:39 +00:00 |
|
kr.angelov
|
1376df457d
|
started an official API to the C runtime
|
2012-12-12 11:25:58 +00:00 |
|