kr.angelov
|
51ae8bbce1
|
the Bulgarian phrasebook works again
|
2013-05-30 13:16:35 +00:00 |
|
kr.angelov
|
739b10f2a8
|
a simple refactoring in the Python runtime
|
2013-05-29 11:02:18 +00:00 |
|
kr.angelov
|
43bffd3f7d
|
readPGF in the Python runtime now throws "No such file or directory" exception if the grammar is missing
|
2013-05-29 10:49:56 +00:00 |
|
kr.angelov
|
bae05df3b0
|
added a test class for the Java API plus a small refinement in the implementation for the binding
|
2013-05-28 13:32:32 +00:00 |
|
kr.angelov
|
3147e16453
|
an initial skeleton for building a Java binding to the C runtime
|
2013-05-28 12:59:19 +00:00 |
|
kr.angelov
|
b7cbee7940
|
fix the encoding problem with unicode literals in the Python binding
|
2013-05-21 10:53:20 +00:00 |
|
kr.angelov
|
517b8ff1ee
|
bugfix for bracketedLinearize which was causing crash if the tree cannot be linearized
|
2013-05-07 08:35:33 +00:00 |
|
kr.angelov
|
687b326ed0
|
bug fix in the management of memory pools in the statistical parser
|
2013-05-07 08:30:32 +00:00 |
|
kr.angelov
|
7ba27229b3
|
the statistical parser is now using two memory pools: one for parsing and one for the output trees. This means that the memory for parsing can be released as soon as the needed abstract trees are retrieved, while the trees themselves are retained in the separate output pool
|
2013-05-06 15:28:04 +00:00 |
|
kr.angelov
|
520c2fb59d
|
fix the leftcorner filtering after the addition of word completion
|
2013-05-05 10:30:06 +00:00 |
|
kr.angelov
|
b8d61fcbb2
|
bugfix in the C runtime which was causing an infinite loop while linearizing partial trees
|
2013-05-04 13:32:57 +00:00 |
|
kr.angelov
|
22f44ef61f
|
word completion in the C runtime. The runtime/python/test.py example is now using readline with word completion
|
2013-05-01 06:09:55 +00:00 |
|
kr.angelov
|
85efdf81e7
|
finally the statistical parser is able to return all possible abstract trees
|
2013-04-26 20:44:01 +00:00 |
|
kr.angelov
|
aad274022a
|
a trivial refactoring of the reasoner in the C runtime
|
2013-04-23 06:40:14 +00:00 |
|
kr.angelov
|
6481f448fc
|
bug fix in pgf-translate which was hiding that there are more than one trees per sentence
|
2013-04-22 13:02:43 +00:00 |
|
kr.angelov
|
4394b8a3bc
|
a bit more informative error message in GrammarToPGF
|
2013-04-22 12:14:39 +00:00 |
|
kr.angelov
|
b1b68bf6b4
|
reverse the direction of the arcs in the dependency trees
|
2013-04-21 19:20:08 +00:00 |
|
kr.angelov
|
09c1bd662d
|
the option -old for the vp command is now redundant
|
2013-04-19 11:15:18 +00:00 |
|
kr.angelov
|
4e2044ab99
|
remove the dead code left behind by Peter Ljunglöf in VisualizeTree
|
2013-04-19 11:13:07 +00:00 |
|
kr.angelov
|
a591160b94
|
fix the command options for the vd command in the shell
|
2013-04-19 11:11:57 +00:00 |
|
kr.angelov
|
542dcaa0ec
|
the C runtime and the Python binding now have an API for parser evaluation. The API computes PARSEVAL and Exact Match for a given tree. As a side effect the abstract trees in Python are now compared for equality by value and not by reference
|
2013-04-19 10:57:46 +00:00 |
|
kr.angelov
|
f050609101
|
added API for computing bracketed strings from Python and C
|
2013-04-18 13:37:09 +00:00 |
|
kr.angelov
|
b49b9d459a
|
added a malt_tab format to the vd command in the GF shell
|
2013-04-16 18:22:37 +00:00 |
|
kr.angelov
|
f6d675c34b
|
the generation of dependency trees in the Haskell runtime is now finally working with bracketed strings. This also fixes some errors in the old implementation
|
2013-04-16 13:10:48 +00:00 |
|
kr.angelov
|
2f35964871
|
the compiler now sorts the list of functions per category in probability order. this ensures probability order search in the C runtime
|
2013-04-15 19:58:57 +00:00 |
|
kr.angelov
|
1f91928606
|
now the web service to the robust parser can to translations also
|
2013-04-05 12:22:52 +00:00 |
|
kr.angelov
|
9e741cfe30
|
a bugfix which was causing an infinite loop in the C linearizer for some sentences
|
2013-04-05 09:11:24 +00:00 |
|
kr.angelov
|
a449a240de
|
a very simple linearization for partial abstract trees in the C runtime
|
2013-04-05 08:42:56 +00:00 |
|
kr.angelov
|
74a16273b9
|
added simple script for estimating the coverage on the PennTreebank
|
2013-03-28 09:15:38 +00:00 |
|
kr.angelov
|
17fc938c20
|
added a few more multiword expressions in DictEng and a few words in the abstract syntax are not tagged with their senses. There is a new statistical model too
|
2013-03-27 20:46:42 +00:00 |
|
kr.angelov
|
885a14e64d
|
added the file treebanks/PennTreebank/ParseEngAbs3.probs which is used by the statistical parser for robust chunking
|
2013-03-25 10:28:53 +00:00 |
|
kr.angelov
|
40fb775a79
|
a long list of prepositions from Wikipedia is now imported in DictEng in addition there are a number of small other changes in the dictionary. The statistical model is updated and is now moved to treebanks/PennTreebank/ParseEngAbs.probs
|
2013-03-25 10:24:24 +00:00 |
|
kr.angelov
|
d1866472eb
|
added configuration file which defines the heads for all syntactic functions in ParseEng
|
2013-03-21 13:39:24 +00:00 |
|
kr.angelov
|
c6e4db8f4a
|
the calculation of lexical_prob in the statistical parser doesn't work properly. It should be fixed but for now I just disabled the optimization
|
2013-03-20 12:28:52 +00:00 |
|
kr.angelov
|
2aacbb0c46
|
replace #if with #ifdef when checking for the optional bottom up filtering in the C runtime
|
2013-03-20 10:47:47 +00:00 |
|
kr.angelov
|
770b1af6d9
|
the bottom up filtering in the C runtime is temporary disabled. It takes too much memory and even makes it impossible to load the Finnish and the German parsing grammars.
|
2013-03-19 10:59:44 +00:00 |
|
kr.angelov
|
411d91d410
|
bug fix in the linearizer in the C runtime
|
2013-03-14 12:31:49 +00:00 |
|
kr.angelov
|
d018502fca
|
update the pgf-service tool from the C runtime after the changes in the API
|
2013-03-14 10:37:01 +00:00 |
|
kr.angelov
|
ca3716857c
|
bugfix in the statistical parser
|
2013-03-11 14:47:43 +00:00 |
|
kr.angelov
|
026c198974
|
added gu_buf_flush in seq.c which removes all elements from a buffer
|
2013-02-26 09:48:09 +00:00 |
|
kr.angelov
|
9cb0b580d3
|
the parser in the C runtime should not crash if the start category is not defined
|
2013-02-19 12:08:48 +00:00 |
|
kr.angelov
|
13de2fafb4
|
bugfix in the grammar reader in the C runtime
|
2013-02-19 12:04:10 +00:00 |
|
kr.angelov
|
9940fe392e
|
bugfix in the linearizer for the C runtime
|
2013-02-13 15:39:01 +00:00 |
|
kr.angelov
|
4922ab6cc4
|
now the beam size for the statistical parser can be configured by using the flag beam_size in the top-level concrete module
|
2013-02-12 10:53:13 +00:00 |
|
kr.angelov
|
a4c9d20fc3
|
the statistical parser now uses a baseline lexical estimation of the beam size
|
2013-02-12 09:41:32 +00:00 |
|
kr.angelov
|
6a36ce77ff
|
the class PgfConcr from the Python binding now has a property name which returns the name of the concrete syntax
|
2013-02-11 15:51:26 +00:00 |
|
kr.angelov
|
d124fa9a12
|
refactoring: now all named objects in the C runtime have an explicit name field
|
2013-02-11 14:10:54 +00:00 |
|
kr.angelov
|
90c3304147
|
remove the pgf2yaml tool which was both broken and redundant. The declarations for generic programming from data.c are removed as well
|
2013-02-11 13:51:12 +00:00 |
|
kr.angelov
|
10ef298fa0
|
the grammar reader in the C runtime is completely rewritten and it doesn't use the generic programming API
|
2013-02-11 10:16:58 +00:00 |
|
kr.angelov
|
5e2474e346
|
This patch removes Gregoire's parse_tokens function in the python binding and adds another implementation which builds on the existing API for lexers in the C runtime. Now it is possible to write incremental Lexers in Python
|
2013-02-01 09:29:43 +00:00 |
|
kr.angelov
|
c99ab058ea
|
implement gu_exn_caught in gu/exn.c. It was missing
|
2013-02-01 09:26:30 +00:00 |
|
kr.angelov
|
eda1058441
|
fix the Windows link
|
2013-01-31 15:06:42 +00:00 |
|
kr.angelov
|
e2d0ab8c62
|
added a link to the Windows binary from the download page
|
2013-01-31 15:03:35 +00:00 |
|
kr.angelov
|
84fa796de4
|
bugfix in the reference counting for Python
|
2013-01-29 09:41:12 +00:00 |
|
kr.angelov
|
05cb74d14a
|
the Python binding is in pure C again
|
2013-01-29 09:20:32 +00:00 |
|
kr.angelov
|
b524c5d8b5
|
added an API for composing and decomposing abstract trees from Python
|
2013-01-29 09:07:41 +00:00 |
|
kr.angelov
|
8846648393
|
fixed typos in the python binding: in a few places pgf_ExprType was used instead of pgf_ExprIterType
|
2013-01-29 09:06:23 +00:00 |
|
kr.angelov
|
580e443a5e
|
fix warnings in pgf-parse.c
|
2013-01-08 12:53:49 +00:00 |
|
kr.angelov
|
9b78da5357
|
now the Python binding has an alternative representation for abstract trees which is composed of Python objects. The new representation is not integrated with the core runtime yet
|
2013-01-07 15:11:12 +00:00 |
|
kr.angelov
|
2c169406fc
|
a new reasoner in the C runtime. It supports tabling which makes it decideable for propositional logic. dependent types and high-order types are not supported yet. The generation is still in decreasing probability order
|
2013-01-07 12:50:32 +00:00 |
|
kr.angelov
|
cade578d04
|
bugfix in the strings library from the C runtime
|
2012-12-27 21:18:46 +00:00 |
|
kr.angelov
|
75696808a7
|
bugfix: the linearizer should not generate extra space at the end of the sentence
|
2012-12-19 11:18:34 +00:00 |
|
kr.angelov
|
87360ccc34
|
bugfix for linearization of metavariables at the root of a tree
|
2012-12-19 10:03:05 +00:00 |
|
kr.angelov
|
a28ccc965c
|
rename linearize.{h/c} to linearizer.{h/c} which follows the convention used in parser.c and reasoner.c
|
2012-12-19 09:17:24 +00:00 |
|
kr.angelov
|
490a3f2286
|
a major reimplementation of the linearizer in the C runtime
|
2012-12-19 09:07:05 +00:00 |
|
kr.angelov
|
ff49d21d13
|
fixed accidental bug in pgf-parse.c
|
2012-12-18 15:42:04 +00:00 |
|
kr.angelov
|
403420be2b
|
the C runtime now can read abstract expressions with literals and meta variables
|
2012-12-18 12:29:30 +00:00 |
|
kr.angelov
|
d12c604f9a
|
debugging infrastructure in the reasoner
|
2012-12-14 21:25:00 +00:00 |
|
kr.angelov
|
16a2c38f38
|
bugfix for the reasoner in the C runtime
|
2012-12-14 21:24:17 +00:00 |
|
kr.angelov
|
8ec7ecacca
|
a bit more flexible API for parsing in Python
|
2012-12-14 16:00:52 +00:00 |
|
kr.angelov
|
20aaa4a989
|
The first prototype for exhaustive generation in the C runtime. The trees are always listed in decreasing probability order. There is also an API for generation from Python
|
2012-12-14 15:32:49 +00:00 |
|
kr.angelov
|
f7a5eb0df1
|
bugfix in the lexer from the C runtime. the input sentence doesn't have to terminate with whitespace
|
2012-12-13 16:45:44 +00:00 |
|
kr.angelov
|
0f0b7158c9
|
added simple error handling in the Python test
|
2012-12-13 16:44:39 +00:00 |
|
kr.angelov
|
75c544027b
|
added a simple test for the Python binding
|
2012-12-13 16:19:56 +00:00 |
|
kr.angelov
|
836b953b9d
|
an initial API for parsing and linearization from Python
|
2012-12-13 15:39:07 +00:00 |
|
kr.angelov
|
14e721dda9
|
a top-level API for parsing in the C runtime
|
2012-12-13 14:44:33 +00:00 |
|
kr.angelov
|
68249a11d2
|
bugfix: the outside probability of a PgfItemConts must always be initialized to zero
|
2012-12-13 11:11:45 +00:00 |
|
kr.angelov
|
2dc8236170
|
bugfix: pgf_read_expr no longer requires a semicolon at the end of an abstract expression
|
2012-12-13 11:09:26 +00:00 |
|
kr.angelov
|
0891ef3f0f
|
an initial Python binding to the C runtime
|
2012-12-12 11:29:39 +00:00 |
|
kr.angelov
|
aa13090b66
|
started an official API to the C runtime
|
2012-12-12 11:25:58 +00:00 |
|
kr.angelov
|
5779887f96
|
bugfix for robust parsing with multi-word units
|
2012-12-11 12:57:22 +00:00 |
|
kr.angelov
|
e174f37940
|
added experimental script for chunking in the C runtime
|
2012-12-03 10:07:54 +00:00 |
|
kr.angelov
|
6e3321d712
|
added INSTALL file and updated README file for the C runtime
|
2012-12-03 09:09:08 +00:00 |
|
kr.angelov
|
5e3b23325e
|
remove the duplicated definition of PgfProductionIdx in parser.c
|
2012-11-19 14:16:31 +00:00 |
|
kr.angelov
|
954d7a7ff5
|
bugfix for the building of bottom-up filter in the C runtime
|
2012-11-16 13:27:15 +00:00 |
|
kr.angelov
|
5c52eaf0b7
|
revised heuristic in the statistical parser
|
2012-11-14 12:34:22 +00:00 |
|
kr.angelov
|
468464faca
|
bugfix in the statistical parser
|
2012-11-13 09:48:23 +00:00 |
|
kr.angelov
|
d1044b202a
|
two simple heuristics which speed up the statistical parser more than seven times.
|
2012-11-12 22:17:40 +00:00 |
|
kr.angelov
|
182e366f5d
|
a simple refactoring in the statistical parser
|
2012-11-12 21:48:22 +00:00 |
|
kr.angelov
|
7ad4436502
|
more counters in the profiler for the statistical parser
|
2012-11-12 15:36:21 +00:00 |
|
kr.angelov
|
9b2487243e
|
now we store the state instead of the offset for every continuation in the chart for the statistical parser
|
2012-11-12 14:04:52 +00:00 |
|
kr.angelov
|
c28056c4e5
|
in the statistical parser: move the outside probability from the parse items to their continuation. this makes the value slot shared between many items
|
2012-11-12 13:43:43 +00:00 |
|
kr.angelov
|
56f3ff8202
|
small refactoring in the C runtime
|
2012-11-12 13:05:35 +00:00 |
|
kr.angelov
|
cce22a7f7a
|
use size_t consistently as the type for constituent indices in the C runtime
|
2012-11-12 12:51:27 +00:00 |
|
kr.angelov
|
6784a4c76e
|
implemented gu_map_count in runtime/c/gu/map.c
|
2012-11-12 12:42:19 +00:00 |
|
kr.angelov
|
c679b08b38
|
use prob_t instead of float in a few places
|
2012-10-29 08:52:56 +00:00 |
|
kr.angelov
|
118333eee8
|
forgot to add one #ifdef
|
2012-10-25 18:37:22 +00:00 |
|
kr.angelov
|
d185938952
|
a major refactoring in the robust parser: bottom-up filtering and garbage collection for the chart
|
2012-10-25 14:42:53 +00:00 |
|
kr.angelov
|
93e3356d02
|
add teyjus/simulator/builtins/builtins.h
|
2012-10-11 11:10:17 +00:00 |
|
kr.angelov
|
b22075e15a
|
added the forgoten libteyjus.pc.in file in the C runtime
|
2012-10-11 04:22:38 +00:00 |
|